A team of researchers has released Loopie, a series of Mixture-of-Experts models that challenges conventional wisdom about scaling Transformer architectures.

The Loopie series includes two models: a 20 billion parameter model with 2 billion active parameters and a 6 billion parameter model with 600 million active parameters. Both use a looped architecture that processes information multiple times through the same network layers.

Looped Transformers have historically struggled with a fundamental scaling challenge. When given N times more compute budget, simply increasing parameter count by factor N typically delivers better performance than looping a smaller model N times.

Breaking the scaling paradigm

The research team, led by Zitian Gao and including Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, and Bryan Dai, conducted extensive ablation studies comparing Loopie against vanilla Transformer baselines.

Their experiments included comparisons with a traditional 30 billion parameter model. Results showed Loopie substantially outperformed standard Transformer architectures when trained with identical compute budgets.

The models employ a Mixture-of-Experts design, activating only a subset of parameters for each forward pass. This approach reduces computational overhead while maintaining model capacity.

Reasoning capabilities

The team developed a novel post-training pipeline that equips Loopie with enhanced reasoning abilities. Early results suggest the models achieve frontier-level performance on mathematical reasoning tasks.

The research addresses a core question in AI scaling: whether architectural innovations can compete with brute-force parameter increases. Loopie's success suggests that clever design choices may offer alternatives to simply building larger models.

The work appears particularly relevant as the industry grapples with diminishing returns from traditional scaling approaches. Training costs and computational requirements have grown exponentially, making architectural efficiency increasingly important.

The paper was submitted to arXiv on July 17, 2026, with a revision published three days later. The researchers have not yet announced plans for public model releases or commercial applications.