Kernel keeps anonymous THPs intact across swap
A 30-patch series introduces PMD-level swap entries so huge anonymous mappings survive swap-out and restore in one fault instead of shattering into PTEs.
Usama Arif has posted version 8 of a large memory-management series that lets the Linux kernel swap out anonymous transparent huge pages (THPs) without first splitting their page-table mappings.
Today, when reclaim pushes a PMD-mapped anonymous THP to swap, the kernel breaks the huge mapping into hundreds of ordinary PTE-level swap entries. Swap-in then costs one fault per page and leaves a scatter of small mappings until khugepaged collapses the range again. The new work adds a compact PMD-level swap entry that encodes a contiguous run of swap slots, so the huge mapping can round-trip intact. On fault, a dedicated PMD swap-in path restores the single large mapping directly when conditions allow.
The design is deliberately conservative. Swap accounting stays per-slot. If the swap cache has been split, a slot lives in zswap, THP policy has changed, or allocation fails, the entry is split and the existing PTE path takes over. Callers that must inspect or free only part of the range (mincore, partial MADV_FREE, and similar) do the same. Soft-dirty, userfaultfd write-protect, migration-style moves, unmap, and related bookkeeping are taught to treat the new entry like other non-present huge PMDs so tools and features do not silently skip swapped-out THPs.
On a vm-scalability sequential swap benchmark (four pinned workers thrashing a 6 GiB anonymous set on a 4 GiB guest with THP always on and zswap off), Arif reports median aggregate throughput rising from about 584 MiB/s to 2409 MiB/s (roughly 4.1×), wall time falling by about 76%, and major faults dropping by about 87%.
For workloads that already rely on anonymous THPs and hit swap under memory pressure, the change removes a long-standing tax: the huge page no longer has to be rebuilt after every trip to disk.