freenode
Databases & Infrastructure

QEMU preallocate filter could lose acknowledged guest writes

Races in multi-iothread setups left container image data reading back as zeroes after the guest was told the writes were on disk.

Denis V. Lunev has fixed data-loss bugs in QEMU's preallocate block filter after a production guest running a container workload lost 45 MiB across 164 files in an image layer. The guest had written the data and been told those writes were durable; the ranges later read back as zeroes. The damage was scattered in guest offsets but contiguous in the underlying file, consistent with a qcow2 image over a range that had been zeroed out from below.

The filter grows the image ahead of extending writes by zero-filling a large trailing span, and it tracks where real data ends and where those preallocated zeroes begin. That bookkeeping was not safe once a single node is driven from several AioContexts at once, as with multiqueue virtio-blk and iothread queue mapping. Two writers extending the file could both pass a compare-and-store on the reserved length; the smaller end could land last. Both writes were acknowledged, yet the tracked end now sat below the larger write. Nothing raised it again except a later write past that point, and closing the image truncated the file to the stale length, cutting the data off the end.

A second path let a preallocation wipe ranges that already held acknowledged guest data. The filter advanced its data-end marker before issuing the fallocate and trusted the block layer to order later writes after that zeroing. The trust was misplaced: a parked serialising request could be ignored by conflict detection, so guest writes passed straight into the range about to be zeroed. In the production case the file grew in thirteen small unaligned steps over a tenth of a second, then once by 82 MiB: a jump that could only come from a zeroing that started 49 MiB lower, exactly the hollowed range, after the fallocate had already run over thirteen acknowledged writes.

The series takes a mutex across the filter's length updates, then reserves the preallocated span up front so concurrent extenders see a length that already covers the in-flight zeroing instead of waiting on the slow fallocate. Writes that would race into that window are held back by the filter itself until preallocations drain. A separate change moves the zero-start mark past any non-zero write that reaches into the preallocated region, so a later write-zeroes request is not skipped as already zero and cannot leave stale guest data behind. qemu-io and blkdebug gain multi-iothread submission and breakpoints in the untracked write window so the races can be hit from tests rather than only from production virtio-blk.

Lunev notes that the storage trace cannot decide which of the two failure modes struck the original guest; both are closed.