Real token volume · Chinese models 46% share · Usage ≠ quality · App layer secrets · August outlook
Still picking models from last year's benchmark charts? OpenRouter is the largest neutral model-routing platform — its leaderboard ranks by real paid token volume, not marketing. This article uses data through July 25, 2026 to decode three July shifts: Xiaomi Mimo V2.5 takes the daily crown, Chinese vendors cross 46% share, and capability diverges from popularity. Includes Top 12 model tables, vendor share, app-layer rankings, pricing comparisons, a six-step tiered routing runbook, and August forecasts — cross-linked with our June rankings deep dive and OpenRouter API guide.
The official OpenRouter leaderboard (openrouter.ai/rankings) sorts by real paid token traffic — it shows what developers actually use, not which model is smartest. Before citing the board for procurement, recognize these five blind spots.
Treating volume as quality: Cheap, fast models behind high-traffic apps can dominate the board without being stronger at complex reasoning. The spend-share picture on hard tasks looks completely different.
Chasing the daily winner only: Rankings swing day to day (Mimo V2.5 and Claude Opus 4.8 swap places easily). Track vendor-level share and 30-day totals alongside daily leaders.
Ignoring the price lever: DeepSeek V4 Flash inputs run $0.05–0.14/M versus GPT-5.5 at ~$5/M — a roughly 35× gap. Volume rankings are largely price math.
Model layer only, no app layer: Model rankings tell you which brain is popular; the Apps board reveals what those brains are doing — coding agents and roleplay consume a huge share of traffic.
Overlooking security and governance: This week's OpenAI sandbox-escape incident is reshaping enterprise selection — vendor security reputation is now a first-class variable.
Once you see past these traps, the July tables below read correctly: a traffic thermometer, not a capability verdict.
As of July 25, daily token volume Top 12 and vendor share (verify live numbers at openrouter.ai/rankings before you publish internally).
| Rank | Model | Vendor | Daily tokens |
|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B |
| 3 | Hy3 | Tencent | 590B |
| 4 | Nemotron 3 Ultra | NVIDIA | 428.6B |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B |
| 6 | GLM 5.2 | Z.ai | 316.7B |
| 7 | MiniMax M3 | MiniMax | 262.5B |
| 8 | Step 3.7 Flash | StepFun | 204.8B |
| 9 | Kimi K3 | Moonshot AI | 157.6B (new entry) |
| 10 | Ling 3.0 Flash | Ant InclusionAI | 128.3B |
| 11 | Gemini 3 Flash | 106.3B | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B |
| Vendor | Region | Share (approx.) |
|---|---|---|
| DeepSeek | China | 16%–18% (most stable #1 single vendor) |
| Xiaomi | China | 8%–18% (Mimo V2.5 surge, highest volatility) |
| Anthropic | US | 10%–15% |
| Tencent | China | 8%–13% |
| US | 8%–13% | |
| Chinese vendors combined ~46% (under 2% one year ago); US big three combined ~30%–36% (about 70% one year ago) | ||
DeepSeek remains the most stable vendor-share champion, but the "model of the month" keeps rotating — MiniMax M2.5, MiMo-V2-Pro, now Mimo V2.5. Do not read the board as "who is #1 today" alone.
| Model | Input/M | Output/M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | $0.05–0.14 | $0.24–0.28 | Best price-performance; agentic coding default |
| GLM 5.2 | $0.45 | $3.31 | Open-weight planning quality closest to Opus tier |
| Kimi K3 | ~$3 | ~$15 | 1.4TB open weights, closed-tier capability |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | Closed flagship; FrontierBench v0.1 43.3% |
By task type, spend share splits roughly: general chat 35.7%, agent workflows 30.4%, code 26.5%. On classification and complex reasoning alone: Claude Sonnet 4.6 and Opus 4.7 tie at 13.5% each, GPT-5.5 third at 11.6% — cheap open models barely register. The market is dumbbell-shaped: open weights capture volume; closed frontier holds pricing power on hard tasks.
Whether you are an indie developer or a platform lead, use these six steps to turn OpenRouter data into an executable tiered routing strategy (pair with fallback setup in our OpenRouter API guide).
Bucket your workloads: Split into volume tasks (chat, creative, roleplay, simple coding) and hard tasks (complex reasoning, high-risk agent decisions, strict classification).
Pick volume-tier models: Test DeepSeek V4 Flash (best economics) and GLM 5.2 (open-weight planning closest to Opus quality). Track acceptance rate and error rate.
Pick hard-tier flagships: Route stuck steps to Claude Opus 5 or GPT-5.6. See our Opus 5 launch breakdown for the new price-performance curve.
Configure OpenRouter fallback chains: On 429 or outage, auto-switch backups. Use the models array for priority order — avoid hard-coded retries in app logic.
Build an internal eval set: Do not trust leaderboard rank alone. Test real prompts for acceptance, latency, and bill size; refresh routing monthly against openrouter.ai/rankings.
Add a security scorecard: Weight vendor safety record, sandbox-escape history, and compliance certs. For autonomous agents, tighten permissions and audit logs.
Tip: Average latency to OpenRouter from many regions runs 180–250ms, and invoicing varies by jurisdiction. OpenRouter fits technical sandboxes; evaluate compliant domestic gateways for production.
| Rank | App | Type | Share (approx.) |
|---|---|---|---|
| 1 | Hermes Agent | Personal agent / CLI | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 6–10 | pi, Lemonade, ISEKAI ZERO, Janitor AI, Cline | Agent / roleplay | 1.7%–3.3% |
Cline → Roo Code → Kilo Code is the same code lineage forked three times — the "grandchild" Kilo Code has already passed ancestor Cline. Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, and similar) drive a large slice of open-model traffic. OpenRouter × a16z State of AI data shows creative roleplay contributes more than half of open-model usage — nearly invisible in enterprise coverage.
Note: Rankings move daily — verify Top 10 at openrouter.ai/rankings before you cite them. Compared with our June analysis, the inflection is Mimo V2.5 displacing DeepSeek Flash as daily leader.
Source: OpenRouter official rankings and public mirrors, through 2026-07-25. Safe to paste into internal docs.
OpenRouter solves multi-model unified routing, but agents running on a local laptop (Hermes, Kilo Code, OpenClaw) still hit sleep, OS updates, and killed processes — smart fallback chains go idle when the host drops offline. Pure Linux VPS lacks Apple Silicon and macOS toolchains. For 24/7 coding agents and iOS CI/CD pipelines, VpsMesh Mac Mini cloud rental is usually the better host: dedicated nodes and auditable environments let tiered routing run on stable infrastructure. See Mac Mini M4 rental pricing and the help center for deployment paths.
As of July 25, Xiaomi Mimo V2.5 leads at roughly 1.4 trillion tokens/day, followed by DeepSeek V4 Flash (943.9B/day) and Tencent Hy3 (590B/day). Daily order shifts — confirm at openrouter.ai/rankings.
On a blended 7-day basis, Chinese vendors total ~46% token share; DeepSeek alone is the most stable single vendor at 16%–18%. Compare with our June 61% model-layer figure — vendor vs model dimensions use different cuts; do not equate them directly.
No. Rankings measure token volume; Claude Sonnet 4.6 and Opus 4.7 still lead spend share on hard tasks. Build an internal eval set and use the help center to deploy stable agent hosts.
For coding, start with DeepSeek V4 Flash + GLM 5.2 hybrid routing; escalate hard steps to Opus 5. Use OpenRouter as a sandbox; for 24/7 agents, plan host costs via Mac Mini rental pricing.
Kimi K3 debuted at #9 (157.6B/day) — the fastest riser. Its 1.4TB open weights set a record. Background: Kimi K3 deep review and distillation controversy recap.