Most teams managing LLM inference costs at scale are working through a familiar checklist: quantisation, KV cache tuning, dynamic batching, speculative decoding. These are sensible interventions. But they all treat the model as a fixed computational unit and try to serve it more efficiently. A different class of technique is emerging that exploits the internal structure of certain model architectures to reduce the computation required per token, without retraining and without a separate model. For teams running looped transformer architectures, this distinction matters significantly.
What Looped Transformers Actually Do Differently
Standard transformer architectures stack discrete layers, each with its own parameters. Looped transformers take a different approach: a single shared block is executed repeatedly across multiple recurrent passes, with each pass refining the model's internal representation before a token is emitted. The practical effect is that you can scale effective computational depth at inference time without scaling parameter count.
This design creates an interesting structural property. Each intermediate pass through the shared block produces a hidden state that is, in principle, decodable into a token distribution. Standard decoding ignores every intermediate state and reads only the output of the final pass. That is a significant amount of available signal being discarded.
The architectural families where this applies include models like Huginn and Ouro, which have seen increasing interest for reasoning-intensive workloads where depth of computation matters more than raw parameter size.
How Contrastive Decoding Repurposes Intermediate States
Contrastive decoding is a technique that guides token selection by comparing the output of a stronger model against a weaker one, amplifying predictions where the stronger model is confident and suppressing those where the weaker model is equally confident. In standard settings, this requires a separate auxiliary model, which adds its own inference overhead and complicates deployment.
Looped architectures remove that dependency. Because earlier recurrent passes represent less computation than the final pass, they naturally function as the weaker signal. The final pass is the stronger. No auxiliary model is needed because the architecture already produces the aligned weak-and-strong prediction pairs internally.
Liu et al. (Liu et al., Hugging Face 2026) formalise this as LoopCD, which operates in two modes. LoopCD-Logits reads an earlier pass through the same output layers and uses the resulting logit distribution as the contrastive signal, at the cost of one additional output projection. LoopCD-Hidden operates directly in hidden-state space, bypassing the output projection entirely and adding zero output overhead.
The FLOP Reduction Trade-off in Practice
The headline result from Liu et al. (Hugging Face 2026) is that applying LoopCD at full recurrent depth improves accuracy, and that accuracy improvement then creates headroom to halve the number of recurrent loops while still matching or exceeding the unguided full-depth baseline. The reported forward FLOP reductions range from 22.5% to 48.2% depending on model and benchmark.
It is worth being precise about what that range reflects. The lower end applies where the contrastive signal provides modest guidance and the loop reduction is conservative. The upper end applies where the architecture is well-suited to the technique and the task benefits from iterative refinement, such as code generation or mathematical reasoning. HumanEval pass@1 for Huginn improved from 22.56% to 31.71% under LoopCD-Hidden, and AIME 2024 pass@1 for Ouro-2.6B-Thinking improved from 61.88% to 73.33% under LoopCD-Logits.
These are not marginal improvements. But they are also not architecture-agnostic. The technique is structurally dependent on the recurrent loop design and does not transfer to standard stacked-layer transformers.
Where This Approach Fails and What to Watch
The failure modes are worth stating plainly before any infrastructure decision is made. First, the technique requires that intermediate hidden states be decodable through the same output head used at the final pass. Architectures where the output projection is tightly coupled to the final layer's representation may not satisfy this condition cleanly.
Second, the quality of the contrastive signal depends on the gap between early and late passes being meaningful but not arbitrary. If the model converges too quickly across loops, early-pass representations will be too close to final-pass representations to provide useful guidance. If the model is unstable across loops, the early-pass signal will be noisy rather than informative.
Third, LoopCD-Logits adds one output projection pass per token, which is not zero cost. For latency-sensitive deployments where even a single additional matrix multiplication is significant, LoopCD-Hidden is the more appropriate variant, accepting that hidden-state guidance is a coarser signal.
Evaluating Whether Your Architecture Is Positioned to Benefit
The practical question for an engineering leader is not whether this technique is theoretically interesting but whether your current or planned model architecture can support it. The relevant conditions are: the model must use a shared block executed across multiple recurrent passes, intermediate hidden states must be decodable through the final output head, and the task profile must include workloads where iterative refinement drives accuracy, such as reasoning, code generation, or structured output tasks.
If your deployment is built on a standard stacked transformer, this technique does not apply. If you are evaluating looped architectures for a future deployment or already running one, the inference cost case is worth modelling carefully. A 40% FLOP reduction at the forward pass level translates directly to throughput and cost per token, which at enterprise scale compounds quickly.
The broader implication is that inference optimisation is increasingly architecture-dependent. Decisions made at model selection time now carry downstream consequences for which decoding strategies are available. Teams that treat model architecture and inference infrastructure as separate decisions are likely to find that the available optimisation space is narrower than it needed to be.
Where Vector Labs Fits
We design and deploy production inference systems where decoding strategy and architecture selection are treated as coupled engineering decisions, not afterthoughts. In our memory-wall analysis, we examined how KV cache efficiency, chip architecture, and high-bandwidth memory interact to shape inference economics at scale. If you are making model architecture decisions that will affect inference costs over a multi-year horizon, contact us at vector-labs.ai/contacts.
FAQs
No. LoopCD is a training-free technique that operates at inference time by repurposing intermediate recurrent states that the model already produces during its forward pass. No fine-tuning, distillation, or auxiliary model training is required, which means it can be evaluated against an existing deployment without model changes.
Not directly. The technique depends on the recurrent loop structure specific to looped transformer architectures, where a shared block is executed multiple times per token. Standard stacked transformers do not produce the aligned weak-and-strong intermediate states that the contrastive signal requires. If your current deployment uses a standard architecture, this technique is not applicable without a model change.
LoopCD-Logits adds one output projection pass per token to read the earlier recurrent state through the output head, which introduces a small but non-zero latency cost. LoopCD-Hidden operates in hidden-state space and adds zero output overhead, making it the better choice for latency-sensitive deployments. The trade-off is that hidden-state guidance is a coarser contrastive signal and may deliver smaller accuracy gains on some tasks.
The reported range of 22.5% to 48.2% forward FLOP reduction reflects the savings from halving recurrent loops while matching unguided full-depth accuracy. The actual saving for your deployment depends on your model, task distribution, and which LoopCD variant you use. We recommend benchmarking on a representative sample of your production workload rather than applying the headline figure directly to cost projections.
The most useful point is during model architecture selection, before committing to a specific model family for a long-running deployment. If you are already running a looped transformer in production, the technique can be evaluated as an inference-time intervention without infrastructure changes. If you are running a standard stacked architecture and considering a migration, the availability of techniques like LoopCD is a legitimate factor in the architecture comparison.

