llama.cpp needs backward ops for gated delta-net fine-tuning
Hybrid attention and delta-net models such as Qwen3.5 cannot run through llama_opt_epoch until several missing gradient rules land.
Hybrid models that mix standard attention with a gated delta-net cannot be fine-tuned in llama.cpp today. The training path builds a backward graph and stops on the first operator that has no gradient rule, so runs of Qwen3.5, Qwen3-Next, and related delta-net architectures abort before any weights move.
The missing pieces are the gated delta-net step itself, the join of convolution state into the micro-batch, the SSM convolution, L2 normalization, and a sigmoid used on the gated output. Without those reverse rules, llama_opt_epoch is unusable for the family.
Contributor joelteply reported the gap and pointed maintainers at a CambrianTech fork that already supplies them: a reverse gated delta-net kernel (CPU, CUDA, and Metal) that returns gradients for every input and the initial state, plus backward implementations for the remaining ops. The fork also tightens CONCAT gradient handling so a non-contiguous layout (for example a transposed attention value tensor) is not sliced incorrectly, and makes the abort path name the unsupported operator.
The same work has already been used for LoRA on Qwen3.5-0.8B and a larger 27B hybrid; full-window gradients through the new rules matched a reference graph at high cosine similarity on the smaller model. The request asks upstream to absorb the kernels or equivalent ones so hybrid delta-net fine-tuning works without a private fork, and notes the broader backlog of operators still missing across backends.