Loop transformers: closing the gap between latent and explicit reasoning

Discover LOTUS, the first latent reasoning method that matches explicit in 3B models, reducing latency 2.5x-6.9x.

miércoles, 1 de julio de 2026 • 2 min read • Q2BSTUDIO Team

LOTUS: latent reasoning that matches explicit in 3B models

In the world of natural language processing and language models, one of the most active frontiers is the optimization of internal reasoning. While traditional approaches, such as explicit chain-of-thought, generate intermediate steps token by token — something costly in latency — an alternative known as latent reasoning emerges. This technique operates on the model's hidden states, replacing sequential decoding with continuous representations that promise greater efficiency. However, until now, latent CoT methods showed inferior performance compared to explicit ones from one billion parameters onward, and that gap widened with scale.

A novel proposal to close that distance comes from loop transformers (or depth recurrence). These are architectures that reuse their weights to increase computational depth without adding parameters, making them ideal for latent reasoning. The key idea is simple: a transformer that processes latent blocks in parallel over multiple iterations, applying a cross-entropy loss on each intermediate step — similar to explicit supervision. In this way, the model learns to reason internally without needing to generate text until the end of the process. The results show that this approach can match the performance of explicit CoT at scales of three billion parameters, reducing the thinking phase latency between 2.5 and 6.9 times, both in compact mathematical expressions and in natural language.

This advancement has direct implications for developing custom applications that require deep reasoning without sacrificing speed. At Q2BSTUDIO we understand that computational efficiency is a critical factor for deploying artificial intelligence in production environments. Our experience in custom software allows us to integrate cutting-edge architectures — such as loop transformers — into solutions that run on aws and azure cloud services, ensuring scalability and optimized costs. Additionally, when it comes to protecting these systems, we offer specialized cybersecurity for AI models, including vulnerability audits in inference pipelines.

The ability to reason without needing to generate intermediate text opens the door to faster and more autonomous AI agents, capable of making complex decisions in real time. For example, in business intelligence service environments, a model that internalizes data analysis can produce insights with power bi without exposing intermediate steps, improving the user experience. At Q2BSTUDIO we develop AI for companies that combines these techniques with agile methodologies, helping our clients transform data into competitive advantages. If you want to explore how to apply latent reasoning in your next project, we invite you to learn about our artificial intelligence for companies solutions. We also offer custom application development that integrates these innovations securely and efficiently.

In short, the gap between latent and explicit reasoning is closing thanks to recurrent designs that make the most of the transformer architecture. This type of progress not only accelerates inference, but also allows more compact and easier-to-interpret models, a step forward towards more practical and sustainable artificial intelligence.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.