STORY RECORD
Laguna S 2.1: coding specialist that fabricates
Poolside's Laguna S 2.1 (118B MoE, 8B active, 1M context, open weights on Hugging Face) self-reports beating trillion-parameter models on coding benchmarks like Terminal-Bench 2.1 (70.2%) and SWE-Bench Multilingual (78.5%).
An independent 160-task eval on r/LocalLLaMA found it the best tool-calling local model tested — 6-level tool chains, zero JSON errors, fastest 100B+ throughput — but also discovered 3 hand-confirmed hard fabrications including an invented financial figure and hallucinated racehorse data, giving it a grounding-under-pressure score of 0.80 vs Qwen 3.5-122B's 0.97.
The model runs on a single DGX Spark, went from kickoff to release in 52 days, and is open under OpenMDW-1.1. All benchmark figures are self-reported using Poolside's own agent harness, and the 40.4% DeepSWE score used a custom harness ineligible for leaderboard comparison.
Poolside's own team acknowledged the model over-thinks, trusts memory over schemas in third-party harnesses, and may produce malformed JSON tool arguments.
The core tension: the same persistence that makes it the strongest open coding agent by tool mechanics also produces confident invention when data runs out, making the 'human review loop' non-negotiable for autonomous deployments.
Why It Matters
This release lands in a moment when the AI efficiency race has shifted from raw parameter count to cost-per-task, and open-weight models from China like Kimi K3 are suddenly topping US labs.
An American company built a model in 8 weeks with only 8B active parameters claiming to match models 200x its size on agentic coding tasks, and it runs on a single desktop appliance you can buy. That alone would be a reset of efficiency expectations.
But the independent discovery of fabrications provides the first real-world stress test of the 'persistence over grounding' tradeoff — the very property that makes Laguna good at long-horizon coding also makes it dangerously confident when it lacks data.
Co-CEO Eiso Kant explicitly framed the release as a stand for open-weight against a future controlled by three or four companies, arguing open models must beat closed equivalents, not just lead their own category.
If these results hold under independent scrutiny, Laguna S 2.1 is simultaneously the strongest argument for efficient open-weight agentic coding AND the clearest warning about the risks of autonomous deployment without human oversight.
The Facts
17Poolside released Laguna S 2.1 on July 21, 2026, a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, supporting up to 1M-token context with thinking and no-thinking modes.
direct primary source · confidence 0.95
The model is open-weight under the OpenMDW-1.1 license, with weights available on Hugging Face and access through OpenRouter and the Poolside API.
direct primary source · confidence 0.95
Laguna S 2.1 runs on a single NVIDIA DGX Spark, with weights shipped in BF16, FP8, INT4, and NVFP4 formats with official GGUF and MLX conversions.
direct primary source · confidence 0.9
Poolside self-reports Laguna S 2.1 scoring 70.2% on Terminal-Bench 2.1 in its own agent harness with thinking enabled.
self reported single source · confidence 0.7
Poolside claims the 70.2% Terminal-Bench 2.1 score exceeds Thinking Machines Inkling (975B, 63.8%) and DeepSeek V4 Pro Max (1.6T, 64%), and approaches same-day Gemini 3.6 Flash (78%).
self reported competitive claim · confidence 0.55
Poolside reports Laguna S 2.1 scoring 78.5% on SWE-bench Multilingual, which Ben Burtenshaw (HuggingFace community MLE) described as 'level with Claude Sonnet 5'.
self reported and anecdotal · confidence 0.5
The 40.4% DeepSWE score was run in Poolside's own 'pool' agent harness, not the official mini-swe-agent harness used by the Datacurve DeepSWE leaderboard, making cross-model comparisons uncertain.
direct primary source · confidence 0.9
The model went from training kickoff to public release in 52 days using Poolside's Model Factory industrialized pipeline.
direct primary source · confidence 0.85
Pi (OpenRouter platform) added Laguna S 2.1 support, highlighting it achieved 78.5% on SWE-bench Multilingual and beats models up to 200x its size.
direct primary source · confidence 0.85
An independent evaluation on r/LocalLLaMA (780K subscribers) by user klinec tested Laguna S 2.1 on a single RTX Pro 6000 Blackwell (96GB) across 160 tasks x 3 runs with deterministic grading, finding 3 hand-confirmed hard fabrications including an invented P&L figure and hallucinated racehorse names, producing a grounding-under-pressure score of 0.80 vs Qwen 3.5-122B's 0.97 (zero inventions across ~240 runs).
independent hands on eval · confidence 0.75
The same independent evaluation found Laguna S 2.1 achieved the best tool-call arg selection (0.89 pass vs Qwen 122B's 0.86), deepest tool chains (6 levels vs 4), and fastest throughput (109 tok/s vs 103 tok/s) among local models tested.
independent hands on eval · confidence 0.75
Poolside's own blog and Pre-training Lead Robert McHardy disclosed known limitations: the model can overthink on hard math with no intermediate effort control, may trust memory of tool interfaces over schemas in third-party harnesses, and may produce incorrectly escaped JSON arrays in tool arguments.
direct primary source · confidence 0.9
Laguna S 2.1 autonomously built a working HTML/CSS rendering engine from scratch in a single 50-minute session of 181 steps with no human intervention, as reported by Poolside.
self reported single source · confidence 0.65
NVIDIA AI's official X account congratulated Poolside on the release, noting the model works with NVIDIA NeMo for customization.
direct primary source · confidence 0.9
Co-CEO Eiso Kant framed the release as a stand for open-weight intelligence against a future controlled by three or four companies, arguing open models must compete with closed equivalents, not merely lead their own category.
direct primary source · confidence 0.8
The independent evaluator reported the model requires max_tokens of 8k+ because at 2048 it burns the entire budget thinking and returns empty, with dflash speculative decoding providing situational speedup (109 -> 271 tok/s on code prompts) but net LOSS under concurrent serving.
independent hands on eval · confidence 0.7
Still Open
1- ContradictionContradiction note is limited to snapshotted evidence.contradicted