PyTorch adds Triton JIT cache hydrate from verified bundles
A precompile import path can stage a checked runtime-cache bundle so GPU kernel launches skip first-use compile and autotune.
PyTorch is adding a precompile path that hydrates Triton's JIT runtime cache from a verified bundle, so workers can launch GPU kernels from known-good entries instead of compiling and benchmarking on first use.
The import validates the bundle, stages cache-key directories under the Triton cache location, renames them into place, and writes a readiness marker last. If the directory is already marked, it checks existing contents against the bundle rather than replacing them. The goal is faster cold starts and more predictable deploys when operators ship a pre-built cache image with the workload.
Review tightened the end-to-end coverage so the consumer records inode and mtime for every cache file before launch and asserts the snapshot is unchanged afterward, catching silent recompiles that a path-only listing would miss. Staging cleanup was adjusted so a failed rmtree cannot hide the import's own error. Leftover staging dirs from a killed import are left alone by design: a PID-based reclaim is unsafe on shared volumes and across container restarts that reuse low PIDs, and a stray staging tree costs at most one bundle's disk.
One concurrency edge remains. The final marker write is not exclusive, so two concurrent imports of different bundles can both return success when keys overlap or interleave, after which a later import that expects the first bundle may reject the directory.