Release notes

HFTKernel 7.2.x release notes.

Browse the 7.2.x kernel line from newest to oldest, with the upstream changes most relevant to predictable CPU execution, stable networking, accurate timekeeping, virtualization, memory behavior, and measurable latency tails.

Public changelog

Release notes

Each entry focuses on upstream changes that support HFTKernel’s primary goals: minimal jitter, controlled tail latency, a stable data path, accurate timekeeping, and measurable behavior under pressure.

HFTKernel 7.2.8

Cleaner memory reclaim, stronger packet capture, and better time diagnostics

Latest 7.2.x entry

Upstream baseline: Linux 7.2.8. Changes since 7.2.7. Upstream release date: September 25, 2026.

More predictable memory reclaim

Linux 7.2.8 improves the split between global filesystem shrinkers and per-memory-cgroup reclaim.

Non-memcg-aware cached-object hooks are kept on the global reclaim path, while memcg-aware filesystem hooks can still participate in per-memcg reclaim. This avoids repeated global work across many memcg and NUMA-node pairs, while restoring useful per-memcg reclaim behavior for paths such as XFS inode slab reclaim.

For HFTKernel, this supports cleaner service isolation and lower background noise on systems that use containers, cgroup v2, or service-level memory boundaries.

More accurate locked-memory accounting

The release corrects locked_vm accounting for MREMAP_DONTUNMAP when locked VMAs are remapped without unmapping the source.

It also makes NR_MLOCK updates safer in paths where munlock(), direct I/O, page-cache folios, and IRQ completion can interact on the same CPU.

For HFTKernel, this strengthens memory accounting around pinned and locked working sets, which are common in latency-sensitive applications and measurement workloads.

Clearer huge-mapping behavior for driver-managed memory

Huge PFNMAP mappings now bypass generic THP tunables where those tunables do not apply. The upstream changelog notes that these mappings are not normal reclaimable THP memory and that incorrect handling caused real-world performance degradation in affected configurations.

For HFTKernel, this is relevant to specialized device mappings, DAX-style mappings, and driver-managed memory, not to the default application THP policy.

Stronger AF_PACKET capture and replay paths

AF_PACKET receives two practical fixes for ring-based packet processing.

TPACKET_V3 now preserves the validated private-area size without truncating it into a smaller internal field. This keeps packet records from overlapping private ring metadata.

The VNET-header error path now clears RX ownership correctly, preventing a failed packet from leaving a ring slot unavailable for the next valid packet.

For HFTKernel, this is directly useful for market data capture, packet replay, latency probes, traffic generators, and network test harnesses built around AF_PACKET.

More robust TCP and socket timestamp behavior

TCP now excludes excessively old ACKs from the fast path and sends them through the normal validation path. This keeps TCP connection state updates aligned with ACK validation.

Socket timestamp setup now takes the socket lock when changing timestamp-related socket flags, while preserving the common lockless case. This strengthens timestamp lifecycle correctness in socket-heavy systems.

For HFTKernel, these changes improve correctness around long-running TCP, TLS, WebSocket, control-plane, and observability connections.

Better PTP and PHC diagnostics

For MACB/GEM, the system timestamps returned around a PHC read now properly bracket the hardware register read. The upstream changelog reports that the previous interval could be unrealistically short compared with the actual ordered register read, biasing midpoint-based synchronization.

For stmmac, phc_index now remains unset when hardware timestamping is supported but no PTP clock has been registered yet. This avoids reporting a valid-looking PHC index for the wrong clock.

For HFTKernel, this improves the integrity of time diagnostics and reduces the risk of misleading PHC discovery in supported NIC configurations.

Stronger RDMA, EFA, and IRQ lifecycle handling

The AWS EFA driver now keeps EQ resources and admin queues alive while their IRQ handlers are registered. Initialization and teardown order are aligned with the resources accessed from interrupt context.

RDMA paths also receive lifecycle and ordering improvements, including local fencing for Intel irdma memory registration and IRQ-safe XArray handling for erdma QP and CQ tables.

For HFTKernel Hardware and cloud-adjacent environments, this strengthens interrupt-driven RDMA lifecycle reliability.

More reliable diagnostics and observability

perf improves cleanup around mediated guest PMU events when modules unload while perf connections are active.

Netlink group metadata is now updated with RCU-safe publication and delayed freeing for lockless readers.

drop_monitor receives safer tracepoint, per-CPU pointer, and timer teardown handling, including timer_shutdown_sync() to prevent rearming during teardown.

For HFTKernel, these changes strengthen the reliability of diagnostic sessions used to investigate packet drops, scheduling noise, and rare tail-latency events.

Included measurement tools

cpujitter

Measures CPU execution jitter, interrupt interference, and latency tails.

memjitter

Measures allocation latency, first-touch behavior, locked working sets, and memory behavior under load.

Use both alongside application-level measurements.

Validate the update with controlled A/B tests using the same toolchain, kernel configuration, CPU and IRQ placement, and workload. Include packet-capture, time-diagnostics, and lifecycle tests for the features used in your deployment.

HFTKernel 7.2.7

More consistent CPU isolation, leaner control paths, and more reliable measurements

Upstream baseline: Linux 7.2.7. Changes since 7.2.6. Upstream release date: September 21, 2026.

CPU isolation and scheduler bookkeeping

Cgroup v2 cpuset management preserves boot-isolated CPUs in its isolation accounting when an isolated partition is released or changes state. This keeps cpuset.cpus.isolated aligned with boot-time domain isolation and supports reliable placement automation across service lifecycles.

EEVDF now maintains its minimum virtual-runtime and minimum/maximum slice metadata consistently during runqueue-tree updates, including initialization of max_slice before enqueue. These refinements support fair-class scheduling for application and housekeeping threads.

Average-idle accounting is updated only when a valid idle timestamp is available, keeping scheduler bookkeeping tied to real idle intervals.

Less unnecessary RT and deadline migration work

RT and deadline schedulers skip migration-disabled tasks when selecting a task to push to another CPU. They can select a movable candidate or return early, avoiding repeated CPU-stopper wakeups that cannot make progress. This improvement applies to the relevant RT/DL scheduling paths.

More efficient deferred console work

Deferred nbcon console output uses lazy IRQ work. When the local scheduler tick is running, the work can be serviced on that tick instead of requesting another immediate interrupt. A CPU with a stopped tick can still require an interrupt, preserving progress in tickless configurations.

More predictable TCP control-packet allocation

TCP active-reset packets use atomic allocation flags, avoiding direct memory reclaim while a socket lock is held. This keeps active-reset allocation lightweight in the affected connection close and disconnect paths, including socket teardown used by network-storage stacks.

More accurate io_uring receive-buffer accounting

Receive operations using MSG_TRUNC account for the bytes actually copied into provided buffers, while preserving the requested full-length result. Incremental buffer consumption stays aligned with the available buffer region, supporting reliable buffer reuse in applications using this mode.

Consistent hardware timestamp diagnostics

The mlx5 driver preserves accumulated hardware timestamp statistics across changes to channel counts, traffic classes, and DMA or port timestamping modes. Persistent statistics also allow the read path to avoid the previous state mutex. This strengthens timestamping telemetry across NIC reconfiguration.

More reliable network-driver lifecycles

Azure MANA: XDP reconfiguration retains the previous valid program state when receive-buffer preallocation cannot complete.

Intel IDPF: dynamic interrupt-moderation work is stopped before queue vectors are released. VLAN headers are also accounted for when parsing hardware-coalesced receive packets.

GRO and forwarding: hardware-aggregated TCP packets are kept compatible with fraglist GRO processing and subsequent segmentation. This is relevant to Linux forwarding configurations combining these offloads.

More dependable performance observability

Intel PEBS sampling coordinates PMU shutdown and shared-buffer draining to prevent overlapping drain operations. This improves the reliability of hardware-assisted performance sampling on supported Intel systems.

Tracing histograms retain percentage and graph modifiers for values. Histogram setup selects and validates the trace clock before publishing the trigger, strengthening repeatable diagnostic-session setup.

Stronger memory and I/O lifecycle management

SLUB coordinates freelist returns with partial-list updates under the appropriate node lock, maintaining consistent allocator state during concurrent operations.

On x86, page tables created when splitting large kernel mappings use the standard kernel page-table allocation path. This keeps construction, accounting, deferred freeing, and IOTLB invalidation coordinated.

For asynchronous io_uring writes, filesystem write accounting is completed from the actual I/O completion callback rather than waiting for later task-work processing. This improves coordination with filesystem freeze and shutdown workflows.

Included measurement tools

cpujitter

Measures CPU execution jitter and latency-tail behavior.

memjitter

Measures memory allocation, first-touch, working-set, and memory-pressure behavior.

Use both alongside application-level measurements.

Validate the update with controlled A/B tests using the same toolchain, kernel configuration, CPU and IRQ placement, and workload. Include steady-state measurements and lifecycle tests for the features used in your deployment.

HFTKernel 7.2.6

More consistent timekeeping, leaner resource management, and stronger network lifecycles

Upstream baseline: Linux 7.2.6. Changes since 7.2.5. Upstream release date: September 14, 2026.

More consistent system timekeeping

Timekeeping now accounts for the continuity-preserving timestamp adjustment in ntp_error when the clocksource frequency multiplier changes. This keeps the clock-discipline error calculation aligned with the adjustment, particularly during larger frequency corrections from an external reference.

More efficient CPU and memory placement management

In cgroup v2, cpuset hotplug handling avoids redundant task updates when effective CPU and NUMA-node masks are inherited from a parent. This reduces unnecessary placement-management work during the affected topology updates.

Less unnecessary memory reclaim work

When a memory cgroup is removed, LRU size accounting moves to its parent instead of remaining duplicated in the child. Traditional LRU reclaim avoids repeatedly scanning emptied lists, while MGLRU benefits from more accurate shadow-node budget accounting. This improves resource bookkeeping during service restarts, container turnover, and memory pressure.

More consistent FQ packet pacing

FQ now compensates only for positive scheduling drift when calculating pacing delays. With pacing offload enabled, early dequeue no longer shortens the next interval in the affected rate-limited and non-EDT paths. Runtime validation of quantum and initial_quantum also keeps FQ configuration within supported bounds.

More reliable AF_XDP and DMA-buffer management

The virtio-net driver now resizes its AF_XDP buffer array together with the RX ring. This keeps buffer capacity aligned with queue configuration when an AF_XDP application is attached.

The shared page-pool code also coordinates DMA unmapping and buffer-metadata cleanup more carefully during concurrent teardown and page returns, strengthening network-buffer lifecycle management.

Stronger hardware and virtual-network configuration

Intel E810 PTP: the ICE driver can fall back to the sideband queue when the low-latency PHY timer interface times out. This improves PTP reinitialization after firmware updates and related resets while retaining the normal low-latency path.

Azure MANA: MSI-X allocation respects the device's actual vector-table capacity, improving interrupt-resource setup on very large VMs.

NVIDIA mlx5: E-Switch vport state changes preserve the configured maximum TX speed, keeping virtual-port configuration consistent across state transitions.

More dependable latency diagnostics

High-resolution timer diagnostics now count interrupt retries that recover successfully, improving the accuracy of nr_retries in /proc/timer_list.

Intel Processor Trace receives more reliable stop/start handling. The updated perf sched tool requires relevant trace samples before presenting a latency table, making incomplete recordings explicit.

More complete I/O accounting across CPU hotplug

Block-layer statistics collect and reset samples from all possible CPUs, including CPUs taken offline during a sampling window. This preserves pending latency samples for the I/O control paths that use these statistics.

Included measurement tools

cpujitter

Measures CPU execution jitter, interrupt interference, and latency tails.

memjitter

Measures allocation latency, first-touch behavior, locked working sets, and memory behavior under load.

HFTKernel 7.2.6 delivers more consistent timekeeping, leaner resource management, and stronger network lifecycles.

HFTKernel 7.2.5

More Predictable Memory, Synchronization, and Latency Diagnostics

Based on Linux 7.2.5. Changes relative to Linux 7.2.4.

More reliable HugeTLB and NUMA control

Gigantic HugeTLB allocation through CMA now validates the preferred NUMA node against the allowed nodemask, falls back to current cpuset limits when no explicit mask is supplied, and can retry after a concurrent cpuset change. This improves placement control for large preallocated buffers, including 1 GiB pages on x86 when using this allocation path.

HugeTLB pool management preserves max_huge_pages when freeing surplus pages. Migrated pages retain their migratable state and active-list membership even within the same NUMA node. Cgroup counters are initialized consistently regardless of CONFIG_DEBUG_VM, including unlimited limits reported as max.

More consistent priority-inheritance synchronization

Dedicated PI-futex scheduling helpers bracket rt_mutex_wait_proxy_lock() without calling sched_submit_work() in the proxy-lock path. This improves synchronization bookkeeping for applications using priority-inheritance futexes.

On PREEMPT_RT, PI-requeue completion avoids a redundant rcuwait_wake_up() in Q_REQUEUE_PI_LOCKED, keeping lock handoff and waiter wakeups consistent. This second change is specific to PREEMPT_RT.

More reliable BPF and ftrace diagnostics

bpf_get_stackid() protects callchain collection and access to the returned per-CPU buffer from preemption. Concurrent ftrace_ops initialization is synchronized while retaining the fast path for initialized objects.

Function-filter operations and event-filter/trigger reads hold the required trace-instance references. These changes strengthen automated profiling and investigation of rare latency outliers on PREEMPT systems.

More accurate hardware profiling

Intel PMU AnyThread availability follows the relevant CPUID capability. Intel LBR user-only filtering checks both branch addresses, keeping samples within the requested scope.

The extended syscall-argument BPF program in perf trace uses bpf_for, making verifier loop termination checks more robust across Clang changes. This tooling improvement requires an updated perf build.

Cleaner one-time socket initialization

DO_ONCE_SLEEPABLE() defers static-key disabling to a system workqueue, outside caller-held locks. This separates global static-key updates from socket locking in paths such as __inet_hash_connect(). The change applies to one-time initialization rather than per-packet processing.

Stronger KVM memory and timer handling

Shadow-MMU lockless aging reuses the reverse-mapping value captured during the lock check. TDP-MMU Accessed-bit updates use LOCK CMPXCHG to remain consistent with concurrent SPTE changes.

KVM-emulated Hyper-V synthetic timers validate 100-nanosecond conversions and cap deadlines at KTIME_MAX, keeping distant expirations in the future. These changes apply to the KVM host kernel. For a cloud guest, the corresponding benefit depends on the provider's host update.

More robust storage and optional memory policies

NVMe multipath coordinates namespace teardown with SRCU readers. NVMe/TCP checks incoming C2HData against command direction early, strengthening the affected network-storage paths.

DAMOS prioritization uses a rolling access sum for more consistent scheme decisions. With tmpfs THP enabled, MADV_HUGEPAGE passes updated VMA flags to khugepaged. These improvements apply to configurations using the respective subsystems and policies.

Included measurement tools

cpujitter

Measures CPU execution jitter, interrupt interference, and latency tails.

memjitter

Measures allocation latency, first-touch behavior, locked working sets, and memory behavior under load.

Validate the workload with an A/B comparison of 7.2.4 and 7.2.5, keeping the toolchain, kernel configuration, CPU/IRQ placement, and workload unchanged.

HFTKernel 7.2.5 strengthens huge-page placement, priority-inheritance synchronization, and latency diagnostics, with more consistent memory and timer handling for KVM hosts.

HFTKernel 7.2.4

Less Unnecessary Kernel Work. More Predictable Infrastructure.

Based on Linux 7.2.4. Changes relative to Linux 7.2.3.

More efficient long-term memory pinning

GUP reference accounting for pinned pages has been corrected. The affected long-term memory-pinning path can now avoid unnecessary local and global drains of internal LRU buffers.

For HFTKernel, this reduces unnecessary cross-CPU work while preparing pinned buffers and supports the goal of minimizing system noise around dedicated trading cores.

Better-bounded filesystem reclaim

JBD2 now counts every checkpoint buffer examined against the scan budget, including busy buffers. It also checks whether rescheduling is needed when skipping busy buffers.

This bounds the work performed in a single shrinker pass and allows the journal lock to be released promptly when rescheduling is requested. On ext4 nodes, the change is useful when filesystem activity coincides with memory pressure.

More coordinated memory management

Reclaim and page-migration loops now explicitly report quiescent states to RCU Tasks. This helps the corresponding grace periods complete even during prolonged memory processing on PREEMPT kernels. The migration change is especially relevant to KVM hosts that use MMU notifier callbacks.

For configurations using vm.defrag_mode=1, non-movable allocations are routed through the appropriate pageblock reclaim and compaction path. Whole-pageblock compaction also uses more suitable memory-selection rules, allowing the allocator to make better progress in this mode.

SLUB refills reusable sheaves without consuming pfmemalloc emergency reserves, preserving those reserves for their intended kernel uses.

More consistent network TX queue handling

The common networking stack improves trans_start initialization when TX queues are activated. Each affected queue receives the correct watchdog reference time, making queue-state monitoring more reliable during interface activation.

The Broadcom bnxt_en driver now rings the doorbell for packets already prepared even when the final packet in a burst fails linearization. The device is notified about queued TX descriptors without waiting for another transmit call.

More reliable PCI interrupt delivery on Hyper-V

During CPU hot-unplug, a pending PCI MSI interrupt can now be resent to its new target CPU through the parent interrupt domain.

For HFTKernel Cloud on Hyper-V, this improves PCI-device reliability when the set of online vCPUs changes. The change specifically addresses interrupt retargeting when a CPU is taken offline.

Lower filesystem I/O overhead

For small buffered overwrites in large folios that are already up to date, write processing stops after the requested range instead of walking the remaining buffer heads. This reduces unnecessary CPU work, including in the affected ext4 write path.

In dm-io, DM_IO_BIO requests clone the original bio instead of rebuilding the page vector. The original iterator and alignment are preserved, simplifying I/O handling in Device Mapper configurations that use this path.

More reliable latency and jitter diagnostics

BPF callchain collection and copying are protected from preemption while accessing the shared per-CPU buffer. This improves stack-collection reliability on PREEMPT kernels.

Reads from trace_pipe are synchronized with tracing sub-buffer resizing, strengthening kernel-event observability while tracing is being reconfigured.

Stricter auxiliary-clock handling

__do_adjtimex() now checks whether an auxiliary-clock read succeeded. When a reading is unavailable, it returns -ENODEV instead of using a time value that was not obtained.

This improves auxiliary-clock state handling in configurations that use this timekeeping API.

More reliable Intel Speed Select configuration

On supported Intel platforms, SST-CP parameter validation is strengthened for frequency ranges, proportional priority, CPU identifiers, and CLOS identifiers.

For HFTKernel Hardware, this improves control over hardware performance settings when configuring compatible systems.

Included measurement tools

cpujitter

Measures CPU execution jitter and the tails of the latency distribution.

memjitter

Analyzes memory latency, allocation behavior, first-touch behavior, and locked working sets.

Evaluate the impact on the target workload through reproducible A/B tests, network metrics, and application latency distributions.

HFTKernel 7.2.4 focuses on reducing unnecessary kernel work, making memory and device servicing more predictable, and providing a reliable foundation for measurement.

HFTKernel 7.2.3

More Predictable Packet Processing in Cloud and Hardware Deployments

HFTKernel has been updated to Linux Kernel 7.2.3.

HFTKernel 7.2.3 includes the Linux 7.2.3 stable-branch changes most relevant to HFT, market data, execution systems, and latency-sensitive infrastructure.

The primary focus of this release is a more reliable packet lifecycle, predictable lockless TX behavior, correct TCP operation on asymmetric networks, and stronger virtualized infrastructure paths.

AF_PACKET and packet capture

The lifecycle of AF_PACKET TX_RING backed by vmalloc memory has been improved.

TX ring memory now remains alive until every associated skb has completed. Timestamp updates and TP_STATUS_AVAILABLE are completed before the pending reference is released, without adding another lock to the TX-completion hot path.

This is especially useful for HFTKernel systems used for:

  • Market data capture
  • Packet replay
  • Latency probes
  • Traffic generators
  • Network test harnesses
  • Measurement systems based on AF_PACKET

The change makes TX completion and ring-memory release more consistent without adding synchronization to the main packet-processing path.

More stable VLAN datapath

VLAN interfaces now preserve a stable hard_header_len, inherit the required headroom and tailroom correctly, and maintain proper header alignment for AF_PACKET SOCK_RAW.

This improves the predictability of lockless TX paths when using:

  • Hardware VLAN offload
  • Virtual network interfaces
  • Containers and network namespaces
  • Cloud networking
  • Packet capture through raw sockets

For HFTKernel, this means more consistent packet metadata and buffer layout when VLAN offload state changes.

Correct TCP MSS on asymmetric networks

Linux 7.2.3 calculates the advertised TCP MSS from the configured interface or route MTU rather than a previously discovered outgoing-path PMTU.

This is particularly relevant to asymmetric cloud configurations, DSR load balancers, and overlay networks where outbound and inbound traffic can traverse paths with different MTU values.

The change avoids a smaller outbound overlay MTU unnecessarily constraining inbound TCP segment size for the lifetime of a connection.

Route-derived MSS is also clamped to TCP_MIN_MSS, making TCP receive-window initialization more robust when route metrics contain unusual values.

For HFTKernel, this improves the predictability of long-lived TCP, TLS, and WebSocket connections used by market data and execution systems.

Clean transition between AVX2 and SSE

The nftables PIPAPO AVX2 lookup implementation now executes vzeroupper before returning to code that may use SSE instructions.

Clearing the upper parts of the YMM registers prevents the documented AVX-to-SSE transition penalty in subsequent SSE code.

This applies to systems where nftables and AVX2 PIPAPO lookup remain active on infrastructure or housekeeping cores. For HFTKernel, it removes another small but avoidable source of unpredictable CPU overhead.

More predictable XFRM and IPsec control plane

XFRM NAT keepalive states are now processed in fixed-size batches, bounding temporary memory usage independently of the total state count.

Locking was also reworked so reference collection and state processing occur in separate phases.

For cloud infrastructure, this provides:

  • Bounded temporary memory usage
  • More robust processing of large XFRM state sets
  • More predictable IPsec control-plane behavior
  • More consistent interaction with softirq and state-teardown paths

These changes are outside the main trading hot path, but they strengthen the operational resilience of cloud nodes and encrypted network overlays.

Stronger TLS zero-copy offload

TLS device offload now handles the fragment limit of an open TLS record more reliably.

When the fragment array becomes full, the record is pushed even when the application uses MSG_MORE and zero-copy splice() paths.

This strengthens:

  • NIC TLS offload
  • TLS_TX_ZEROCOPY_RO
  • Zero-copy send paths
  • High-load encrypted TCP streams

For HFTKernel, the change is useful on infrastructure nodes where TLS termination or kernel TLS offload is performed directly by the network adapter.

KVM and AMD SEV/SNP

KVM improves AMD SEV and SEV-SNP handling:

  • SEV-specific guest_memfd hooks are connected only when SEV support is enabled
  • Static-call infrastructure can eliminate an unnecessary CALL+RET when SEV is not used
  • Guest-controlled VMSA tracking is more robust
  • A vCPU correctly returns to runnable state after AP_CREATE
  • File-backed memory support for encrypted guests is restored
  • Temporary buffers for encryption operations are allocated as full pages

These changes primarily matter when HFTKernel is used as a KVM host or with AMD confidential virtualization. Their direct impact inside an ordinary cloud VM depends on the provider host kernel.

NVMe/TCP and Soft RoCE

Use of page_frag_cache during concurrent NVMe/TCP namespace creation is now serialized, improving PDU-buffer lifecycle during parallel storage initialization.

Soft RoCE now updates responder resources more consistently when IB_QP_MAX_DEST_RD_ATOMIC changes: replacement arrays are installed only after the responder task has stopped.

These changes are useful for HFT infrastructure that uses network storage, software RDMA, lab environments, or virtualized RDMA validation. They do not change the primary hardware-RDMA fast path.

eBPF validation

The BPF sockmap selftest is corrected for packets with a non-linear skb.

The parser remains read-only, while the required packet data is explicitly linearized in the verdict program before direct access.

This improves regression testing for BPF sockmap paths and strengthens the HFTKernel build and qualification pipeline.

Impact assessment

Linux 7.2.3 does not publish new median, p99, p99.9, or p99.99 latency measurements for these changes.

The practical effect of the release should therefore be verified with controlled A/B measurements on the target workload.

HFTKernel 7.2.3 strengthens the Linux networking paths where packet lifecycle, buffer metadata, TCP path selection, and background infrastructure operations can affect the stability and repeatability of latency-sensitive systems.

Included measurement tools

HFTKernel 7.2.3 also includes cpujitter and memjitter as native, release-aligned measurement tools.

cpujitter

measures CPU execution jitter, IRQ and softirq interference, and extreme tail latency.

memjitter

measures allocation latency, first-touch behavior, locked working sets, and memory latency under load.

The runtime and measurement tools remain aligned to the same immutable Cloud and Hardware release matrix.

HFTKernel 7.2.3 strengthens packet lifecycle, VLAN metadata, TCP MSS selection, TLS offload, virtualized I/O, encrypted KVM, storage, and validation paths for a more stable and predictable HFT platform.

HFTKernel 7.2.2

More Reliable GSO Fragment Handling in Cloud Networking

HFTKernel has been updated to Linux Kernel 7.2.2.

This focused release strengthens the handling of fragmented IPv4 and IPv6 packets in virtualized and userspace-mediated network configurations.

Network datapath

When an IP fragment enters the shared reassembly queue, the kernel now resets its GSO state with skb_gso_reset().

  • GSO state is no longer carried from individual fragments into the reassembled packet.
  • The reassembled packet reaches software segmentation with the correct state.
  • The shared fix covers IPv4, IPv6, Netfilter connection tracking, and 6LoWPAN reassembly.
  • The interaction between fragment reassembly, GSO, and skb_segment() is more robust.

The change is implemented in the common inet_frag_queue_insert() path and touches only one kernel file.

Practical value for HFTKernel

The update is especially relevant to systems that use:

  • TUN/TAP
  • AF_PACKET with PACKET_VNET_HDR
  • Virtio-net and KVM
  • Cloud virtual machines
  • Network namespaces
  • Packet capture and replay
  • Traffic generators and network test harnesses
  • Virtualized market data and execution gateways

For HFTKernel, this means a more resilient network datapath and more reliable handling of rare combinations of IP fragmentation and GSO metadata.

Upstream validation

Upstream testing covered four different geometries of fragmented IPv4 and IPv6 packets. After the fix, every test datagram was delivered completely and without kernel messages.

Impact assessment

This improvement targets operational resilience rather than acceleration of the ordinary network fast path.

Upstream does not publish new throughput, median-latency, or tail-latency measurements for Linux 7.2.2. The correct HFTKernel formulation is:

HFTKernel 7.2.2 improves the reliability of GSO and fragment-reassembly paths in virtualized and userspace-mediated network configurations.

Native measurement tools included

HFTKernel 7.2.2 also includes cpujitter and memjitter as native measurement tools.

cpujitter

measures CPU execution jitter, IRQ and softirq interference, and extreme tail latency.

memjitter

analyzes allocation latency, first-touch behavior, locked working sets, and memory behavior under load.

HFTKernel 7.2.2 strengthens the network paths used by cloud infrastructure, virtual machines, and traffic capture and replay systems, making the Linux platform for HFT more resilient and predictable.

HFTKernel 7.2

Cleaner CPU Isolation, More Accurate Timing, and Better Tail-Latency Control

HFTKernel has been updated to Linux Kernel 7.2.

This release strengthens the kernel paths that matter most to HFT, market data, execution gateways, and other latency-sensitive infrastructure.

Key improvements

  • Cleaner NO_HZ_FULL CPU isolation with less housekeeping activity reaching latency-critical cores
  • More accurate correlation across PTP, PHC, TSC, KVMclock, and Hyper-V timing paths
  • Cache-aware scheduling and improved last-level-cache locality for related workloads
  • Better futex correctness and targeted membarrier scalability
  • Lower-overhead lockless io_uring task work and stronger zero-copy networking paths
  • NAPI, TCP, and Azure MANA improvements for a more stable network data path
  • NUMA, MGLRU, contiguous-memory, and /proc improvements for more predictable system behavior
  • Lower-overhead IRQ and RTLA observability for finding rare latency events

Native measurement tools included

HFTKernel 7.2 includes cpujitter and memjitter as native packages for every supported Cloud and Hardware target.

cpujitter

measures CPU execution jitter, IRQ and softirq interference, and extreme tail latency.

memjitter

analyzes allocation latency, first-touch behavior, locked working sets, and memory performance under pressure.

Together, HFTKernel, cpujitter, and memjitter provide both the optimized runtime and the measurement tools needed to validate latency-sensitive infrastructure.

HFTKernel 7.2 strengthens CPU isolation, timing accuracy, cache locality, network processing, memory behavior, and latency observability as one measurable platform for HFT.

Release philosophy

Upstream stability, workload-specific policy, and measurable outcomes

HFTKernel starts from an exact stable Linux release, records the toolchain and build identity, applies the selected Cloud or Hardware profile, packages cpujitter and memjitter with the runtime, and validates the result as an immutable release.

Exact baselineVerified stable Linux source and pinned kernel version.
Controlled profileCPU, memory, network, timekeeping, and isolation policy matched to the platform.
Native packagesDEB, RPM, or TGZ bundles for the supported production matrix.
Measured deliverycpujitter and memjitter provide repeatable evidence at the tail.

Deployment planning

Evaluate the release against your CPU, NIC, memory, and timing topology.

Package selection is only the first step. The production result depends on boot parameters, runtime affinity, queue placement, PTP/PHC design, memory policy, and a controlled measurement plan.