OpenRouter Rankings July 2026:
Who's Actually Winning the AI Model Race

Real token volume · Chinese models 46% share · Usage ≠ quality · App layer secrets · August outlook

OpenRouter July 2026 model rankings vendor share chart

Still picking models from last year's benchmark charts? OpenRouter is the largest neutral model-routing platform — its leaderboard ranks by real paid token volume, not marketing. This article uses data through July 25, 2026 to decode three July shifts: Xiaomi Mimo V2.5 takes the daily crown, Chinese vendors cross 46% share, and capability diverges from popularity. Includes Top 12 model tables, vendor share, app-layer rankings, pricing comparisons, a six-step tiered routing runbook, and August forecasts — cross-linked with our June rankings deep dive and OpenRouter API guide.

01

OpenRouter rankings for model selection: five common blind spots

The official OpenRouter leaderboard (openrouter.ai/rankings) sorts by real paid token traffic — it shows what developers actually use, not which model is smartest. Before citing the board for procurement, recognize these five blind spots.

  1. 01

    Treating volume as quality: Cheap, fast models behind high-traffic apps can dominate the board without being stronger at complex reasoning. The spend-share picture on hard tasks looks completely different.

  2. 02

    Chasing the daily winner only: Rankings swing day to day (Mimo V2.5 and Claude Opus 4.8 swap places easily). Track vendor-level share and 30-day totals alongside daily leaders.

  3. 03

    Ignoring the price lever: DeepSeek V4 Flash inputs run $0.05–0.14/M versus GPT-5.5 at ~$5/M — a roughly 35× gap. Volume rankings are largely price math.

  4. 04

    Model layer only, no app layer: Model rankings tell you which brain is popular; the Apps board reveals what those brains are doing — coding agents and roleplay consume a huge share of traffic.

  5. 05

    Overlooking security and governance: This week's OpenAI sandbox-escape incident is reshaping enterprise selection — vendor security reputation is now a first-class variable.

Once you see past these traps, the July tables below read correctly: a traffic thermometer, not a capability verdict.

02

July 2026 OpenRouter model and vendor comparison tables

As of July 25, daily token volume Top 12 and vendor share (verify live numbers at openrouter.ai/rankings before you publish internally).

Top 12 models by daily token volume (July 25)

RankModelVendorDaily tokens
1Mimo V2.5Xiaomi1.4T
2DeepSeek V4 FlashDeepSeek943.9B
3Hy3Tencent590B
4Nemotron 3 UltraNVIDIA428.6B
5DeepSeek V4 ProDeepSeek413.7B
6GLM 5.2Z.ai316.7B
7MiniMax M3MiniMax262.5B
8Step 3.7 FlashStepFun204.8B
9Kimi K3Moonshot AI157.6B (new entry)
10Ling 3.0 FlashAnt InclusionAI128.3B
11Gemini 3 FlashGoogle106.3B
12Claude Sonnet 5Anthropic99.5B

Vendor token share (blended 7-day basis, approximate)

VendorRegionShare (approx.)
DeepSeekChina16%–18% (most stable #1 single vendor)
XiaomiChina8%–18% (Mimo V2.5 surge, highest volatility)
AnthropicUS10%–15%
TencentChina8%–13%
GoogleUS8%–13%
Chinese vendors combined ~46% (under 2% one year ago); US big three combined ~30%–36% (about 70% one year ago)

DeepSeek remains the most stable vendor-share champion, but the "model of the month" keeps rotating — MiniMax M2.5, MiMo-V2-Pro, now Mimo V2.5. Do not read the board as "who is #1 today" alone.

Pricing and positioning: volume board vs hard-task spend share

ModelInput/MOutput/MPositioning
DeepSeek V4 Flash$0.05–0.14$0.24–0.28Best price-performance; agentic coding default
GLM 5.2$0.45$3.31Open-weight planning quality closest to Opus tier
Kimi K3~$3~$151.4TB open weights, closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed flagship; FrontierBench v0.1 43.3%

By task type, spend share splits roughly: general chat 35.7%, agent workflows 30.4%, code 26.5%. On classification and complex reasoning alone: Claude Sonnet 4.6 and Opus 4.7 tie at 13.5% each, GPT-5.5 third at 11.6% — cheap open models barely register. The market is dumbbell-shaped: open weights capture volume; closed frontier holds pricing power on hard tasks.

03

Six-step runbook: tiered model routing by task type

Whether you are an indie developer or a platform lead, use these six steps to turn OpenRouter data into an executable tiered routing strategy (pair with fallback setup in our OpenRouter API guide).

  1. 01

    Bucket your workloads: Split into volume tasks (chat, creative, roleplay, simple coding) and hard tasks (complex reasoning, high-risk agent decisions, strict classification).

  2. 02

    Pick volume-tier models: Test DeepSeek V4 Flash (best economics) and GLM 5.2 (open-weight planning closest to Opus quality). Track acceptance rate and error rate.

  3. 03

    Pick hard-tier flagships: Route stuck steps to Claude Opus 5 or GPT-5.6. See our Opus 5 launch breakdown for the new price-performance curve.

  4. 04

    Configure OpenRouter fallback chains: On 429 or outage, auto-switch backups. Use the models array for priority order — avoid hard-coded retries in app logic.

  5. 05

    Build an internal eval set: Do not trust leaderboard rank alone. Test real prompts for acceptance, latency, and bill size; refresh routing monthly against openrouter.ai/rankings.

  6. 06

    Add a security scorecard: Weight vendor safety record, sandbox-escape history, and compliance certs. For autonomous agents, tighten permissions and audit logs.

i

Tip: Average latency to OpenRouter from many regions runs 180–250ms, and invoicing varies by jurisdiction. OpenRouter fits technical sandboxes; evaluate compliant domestic gateways for production.

04

App-layer secrets: coding agents dominate, plus August predictions

OpenRouter Apps Top 10 (token share, approximate)

RankAppTypeShare (approx.)
1Hermes AgentPersonal agent / CLI~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
6–10pi, Lemonade, ISEKAI ZERO, Janitor AI, ClineAgent / roleplay1.7%–3.3%

Cline → Roo Code → Kilo Code is the same code lineage forked three times — the "grandchild" Kilo Code has already passed ancestor Cline. Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, and similar) drive a large slice of open-model traffic. OpenRouter × a16z State of AI data shows creative roleplay contributes more than half of open-model usage — nearly invisible in enterprise coverage.

Five August predictions

  • Chinese open-source share keeps climbing toward 50% this year unless US vendors cut prices sharply.
  • "Model of the month" keeps rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are still in a price war.
  • Anthropic may ship a cheaper tier to chase volume; Opus 5 is already the fourth flagship in two months (Mythos 5 → Fable 5 → Sonnet 5 → Opus 5).
  • Kimi K3's 1.4TB weights should see community quantization builds in 2–4 weeks — see our Kimi K3 deep review.
  • Security and governance become hard selection criteria: OpenAI sandbox escapes, US "AI kill switch" legislation, and White House pre-release review frameworks are expected to land before August ends.
!

Note: Rankings move daily — verify Top 10 at openrouter.ai/rankings before you cite them. Compared with our June analysis, the inflection is Mimo V2.5 displacing DeepSeek Flash as daily leader.

05

Citable hard data and production environment wrap-up

Source: OpenRouter official rankings and public mirrors, through 2026-07-25. Safe to paste into internal docs.

  • Daily leader: Xiaomi Mimo V2.5 at roughly 1.4T tokens/day — July's largest single-day volume.
  • Share reversal: Chinese vendors combined ~46%, up from under 2% in 12 months — one of the steepest share migrations in AI.
  • Price anchor: DeepSeek V4 Flash vs GPT-5.5 input pricing differs by ~35× — the core variable explaining volume-board structure.
  • App concentration: Hermes Agent alone holds ~45% of platform app tokens; coding agents occupy most Top 10 slots.
  • Quality anchor: Claude Opus 5 scores 43.3% on FrontierBench v0.1, above GPT-5.6 Sol at 37.5%.

OpenRouter solves multi-model unified routing, but agents running on a local laptop (Hermes, Kilo Code, OpenClaw) still hit sleep, OS updates, and killed processes — smart fallback chains go idle when the host drops offline. Pure Linux VPS lacks Apple Silicon and macOS toolchains. For 24/7 coding agents and iOS CI/CD pipelines, VpsMesh Mac Mini cloud rental is usually the better host: dedicated nodes and auditable environments let tiered routing run on stable infrastructure. See Mac Mini M4 rental pricing and the help center for deployment paths.

FAQ

Frequently asked questions

As of July 25, Xiaomi Mimo V2.5 leads at roughly 1.4 trillion tokens/day, followed by DeepSeek V4 Flash (943.9B/day) and Tencent Hy3 (590B/day). Daily order shifts — confirm at openrouter.ai/rankings.

On a blended 7-day basis, Chinese vendors total ~46% token share; DeepSeek alone is the most stable single vendor at 16%–18%. Compare with our June 61% model-layer figure — vendor vs model dimensions use different cuts; do not equate them directly.

No. Rankings measure token volume; Claude Sonnet 4.6 and Opus 4.7 still lead spend share on hard tasks. Build an internal eval set and use the help center to deploy stable agent hosts.

For coding, start with DeepSeek V4 Flash + GLM 5.2 hybrid routing; escalate hard steps to Opus 5. Use OpenRouter as a sandbox; for 24/7 agents, plan host costs via Mac Mini rental pricing.

Kimi K3 debuted at #9 (157.6B/day) — the fastest riser. Its 1.4TB open weights set a record. Background: Kimi K3 deep review and distillation controversy recap.