What happened

OpenRouter, which routes API traffic to hundreds of models, analyzed its request logs from January 1 through June 14, 2026 — more than 457 trillion tokens. The headline: DeepSeek roughly doubled its token share from 9% to 18%, and has been the platform's top model author since mid-May. By early June, Chinese models as a group passed American models in token volume — a sharp reversal from 2025, when US models carried about three-quarters of traffic.

The engine is agentic work. OpenRouter classifies traffic by API key using a seven-signal score (tool-call rate, turn count, gap timing, and more), and found agent requests burn about 15 times more tokens than human-led ones; agentic tokens passed human tokens around February 1. DeepSeek's V4 family, released April 24, was its first entry into that race — and within a month, V4 Flash carried 70% of DeepSeek's agentic token flow.

The economics explain the stampede. V4 Flash is MIT-licensed, a ~284B-parameter mixture-of-experts with ~13B active parameters and a 1M-token context, scoring 79.0% on SWE-bench Verified — within 1.6 points of its much larger V4 Pro sibling, whose 80.6% is the top open-weights score. DeepSeek's first-party API lists Flash at $0.14 in / $0.28 out per million tokens; OpenRouter notes GPT-5.5 runs $5 in / $30 out. The catch: DeepSeek's own endpoint retains and trains on your data. Western hosts that don't (Fireworks, Together, DeepInfra) charge roughly double — still a fraction of frontier pricing.

The field is broader than DeepSeek. OpenRouter's June open-weights roundup puts Z.ai's GLM 5.2 — released mid-June with a 1M-token context window and MIT license — at #1 among open models on Artificial Analysis's Intelligence Index (51, about five points below Claude Fable 5), and effectively level with GPT-5.5 on its real-world agentic benchmark. Z.ai reports the model trails Claude Opus 4.8 by one point on FrontierSWE, under Z.ai's own evaluation setup. Vercel added GLM 5.2 to its AI Gateway on launch day, citing the context-window jump from 200K in GLM 5.1. Xiaomi, MiniMax, and Tencent models also gained share — largely, per OpenRouter, at the expense of Google and OpenAI. Even NVIDIA joined the open-weights push with Nemotron 3 Ultra, the strongest US open entrant.

Why it matters

Open weights give companies real choices about cost, hosting, and data control. OpenRouter describes V4 Flash as a plausible substitute for some frontier-model agent workloads at a fraction of the price, and observes that open models have held a consistent 3–6-month capability gap behind the frontier for over 18 months — the frontier is not accelerating away. For any fixed level of intelligence, the price only drops.

There are limits to the data. OpenRouter's figures describe traffic through OpenRouter, not the whole market, and token share is not customer share or spend — hobbyist usage routes nearly a third of its tokens to DeepSeek, which skews volume. Provider location, data-retention policy, and regulation matter as much as the benchmark chart once sensitive company data enters the prompt.

The fine print

OpenRouter flags that V4 Flash is text-only and reportedly needs more explicit prompting than Anthropic models. Outsiders do not know whether DeepSeek's rock-bottom pricing is subsidized by compute, harvested training data, or both. New GLM 5.2 providers vary in quality, and its reasoning consumes many output tokens. The benchmarks are vendor and OpenRouter claims without independent replication.

The frontier labs are selling intelligence by the token. The open-weight labs are discovering that the fastest way into the enterprise may be through the procurement spreadsheet.