llama.cpp Jinja templates allow exponential DoS via nested loops
Uncapped nested range loops in chat templates can pin llama-server before it finishes starting.
The Jinja chat-template renderer in llama.cpp has no iteration budget and no nesting-depth limit, so a short template of nested {% for i in range(2) %} loops can force an exponential number of iterations and hang the process.
Reporter west-ice-g showed that cost scales as 2^N with nesting depth while the template itself stays only linear in size. A depth-12 template (a few hundred bytes) rendered in under a tenth of a second; depth 20 took tens of seconds; depth 24 still had not finished after three minutes. Each extra four levels of nesting multiplied runtime by roughly sixteen, matching the expected exponential growth.
Chat templates are caller-controlled input. They arrive from a model's tokenizer.chat_template field or from --chat-template / --chat-template-file, and they are applied through the library's template path before llama-server binds its HTTP port. On a server that means a hostile or merely pathological template can pin the rendering thread at startup, a straightforward remote denial of service against any deployment that accepts untrusted templates.
The issue was filed against current llama.cpp development builds and specifically names llama-server among the affected modules. Without a hard cap on loop iterations or nesting depth, any service that renders third-party or model-supplied Jinja chat templates remains exposed to the same blowup.