
Yeltsin in the AI Aisle

The AI model market has segmented along cost, speed, and accuracy dimensions, making older open-weight models competitive long after their release.
OpenRouter data shows that GPT-OSS 120b, a year-old open-source model, still serves 36% of the daily tokens of Anthropic’s newer Claude Opus 4.8. GLM 5.
2 now serves 495 billion tokens per day, outpacing frontier models. Anthropic shipped Opus 5 to contest the space that Moonshot’s Kimi 3 targets, while Poolside launched Laguna S 2.
1 for the US mid-market. The key technical driver is mixture-of-experts (MoE) architecture. The 118-billion-parameter Laguna S 2.
1 activates only 8 billion parameters per token, enabling it to decode at the same speed as a dense 26B model on Apple M5 Max hardware.
Replacing Gemma 4 26b with Laguna in a local agent stack reduced tool-call failure rates from 29.4% to 20.1%, a 7 percentage point gain.
Segmentation shows healthy competition: frontier models still handle the hardest tokens, but lower tiers now achieve higher quality ceilings.


