Radar · 23/07/2026 · models

Laguna S 2.1: 118B open-weight outperforms Claude Fable 5 and costs less than DeepSeek v4 Flash

Poolside AI releases Laguna S 2.1, a 118-billion-parameter open-weight model with MoE architecture (8B active per token), that outperforms Claude Fable 5 on benchmarks and costs less than DeepSeek v4 Flash. Context window up to 1 million tokens, with thinking and no-thinking modes.

Why it matters to you. A team of fewer than 70 researchers produces a model that beats frontier models ten times larger. As we covered on July 20 with Qwen 3.8, the gap between open and proprietary models is narrowing: choosing between them becomes economic, not technical. Kimi K3, Qwen 3.8, and now Laguna S 2.1 converge on the same signal: quality compression is a trend, not an isolated case.

For those using AI in their work, the cost of an open-weight API is no longer the price of a weak model. You can pay less and get results comparable to frontier models on many tasks. The math changes, and it’s worth redoing on your own data instead of staying locked to July benchmarks.

Details

Poolside AI is a neolab founded by Eiso Kant, who spent four years and 12 million dollars building models for code before the market noticed. Today the team has fewer than 70 researchers but has built what Kant calls the “Model Factory”: an end-to-end system that takes a model from pre-training to release in eight weeks, running 10,000-20,000 experiments per month.

Laguna S 2.1 is a Mixture-of-Experts: 118 billion total parameters, but only 8 billion active per generated token. It’s like having a team of specialists where only the relevant expert answers each question, instead of mobilizing the entire department. This keeps inference costs low while maintaining high overall capacity. The context window reaches 1 million tokens and the model offers thinking mode (reasons before responding) and no-thinking mode (responds directly).

On benchmarks, Laguna S 2.1 outperforms Claude Fable 5 and Thinking Machines’ model (Inkling, 975B parameters released July 17), with a model nearly ten times smaller. It costs less than DeepSeek v4 Flash and beats DeepSeek v4 Pro on quality.

Context matters. On July 17, Moonshot AI released Kimi K3, 2.8T parameters, the largest open-weight model ever published. Three days later Alibaba previewed Qwen 3.8, 2.4T parameters. Both showed that open model quality approaches frontier levels. Poolside adds a third data point: you don’t even need to be huge. A well-trained 118B MoE, with the right experimentation pipeline, is enough.

The interesting part of Poolside’s strategy is emphasis on persistence, verification, and backtracking. Kant argues that the ability to attempt, verify, and correct matters more than raw intelligence. It’s the same direction we see in coding agents, where control loop architecture beats single-model power.

Limitations in what we know today. The cited benchmarks come from Poolside’s tech report and Latent.Space coverage from July 23, not independent verified evaluations. The model is open-weight but we don’t yet have adoption data or real-world task testing beyond coding. Poolside raised 500 million dollars, so it has resources to sustain research, but the open model landscape moves so fast that today’s advantage may last weeks.

For those building products on a single model, the question changes. Today the real question is which combination of quality, cost, and license lets you sleep at night. A 118B MoE with 8B active costs little per call, and if weights are open you can self-host it too. The weak point remains verification: as long as numbers come from the lab that built the model, the advice stays the same, test it on your own tasks before believing the rankings.

Type to search across course, playbooks, skills, papers…