Luna -80% · Terra -20% · Sol Fast Mode · Kimi K3 Pressure · Self-Optimized Infrastructure
On July 30, 2026, OpenAI cut API prices for two of its three GPT-5.6 models: Luna dropped 80% to $0.20/$1.20 per million input/output tokens, and Terra fell 20% to $2/$12. Sol kept its price but gained a Fast mode that costs twice as much for up to 2.5x the speed. If you are sizing Agent workflow costs or multi-model API routing, this article delivers the full three-week timeline, tiered pricing and competitor tables, self-optimized infrastructure analysis, controversy notes, and a six-step API selection runbook. OpenAI says part of the savings came from Sol autonomously rewriting its own production GPU code, against a backdrop of competitive pressure from Kimi K3 and DeepSeek. Data as of 2026-07-31
This is not an isolated promotion. It is a fast three-act script—launch, reveal cost-saving engineering, then cut prices—that is unusually quick in OpenAI history:
| Date | Event |
|---|---|
| July 9 | OpenAI launches GPT-5.6: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens |
| July 16 | Moonshot AI releases Kimi K3, 2.8T MoE, priced $3/$15 ($0.30 on cache hits)—nearly half Sol's rate; US tech stocks dipped |
| ~July 27 | Kimi K3 open weights become downloadable, adding self-hosting pressure |
| July 29 | OpenAI details how Sol in Codex rewrote production GPU kernels (Triton, Gluon) and optimized speculative decoding |
| July 30 | Official Luna/Terra price cuts and Sol Fast mode launch; Sam Altman recently called cost "a huge issue" |
| July 31 | Coverage snowballs across CNBC, Reuters-sourced reports, and Chinese outlets |
Five common misconceptions around this repricing:
Treating it as a blanket discount: Sol standard pricing held steady while a more expensive Fast mode was added—tiered pricing, not uniform markdown.
Missing Luna's strategic intent: Luna targets high-volume Agent workloads; -80% is a direct play for price-sensitive customers, not random promotion.
Comparing sticker prices only: Artificial Analysis puts Sol vs Kimi K3 at about $1.04 vs $0.94 per completed task—far closer than per-token rates suggest.
Trusting the 20% cost-cut figure blindly: The "AI rewrote GPU code" narrative comes entirely from OpenAI's blog with no independent audit.
Assuming open weights always mean cheaper: Kimi K3 is roughly 6x pricier than Moonshot's own K2.6—open-weight does not automatically mean lowest price.
| Model | Old (in/out per 1M) | New | Change |
|---|---|---|---|
| GPT-5.6 Luna | $1.00 / $6.00 | $0.20 / $1.20 | -80% |
| GPT-5.6 Terra | $2.50 / $15.00 | $2.00 / $12.00 | -20% |
| GPT-5.6 Sol (Standard) | $5.00 / $30.00 | $5.00 / $30.00 | No change |
| GPT-5.6 Sol (Fast mode) | N/A | $10.00 / $60.00 | 2x standard, up to 2.5x speed |
Note: Sol Fast replaces Priority Processing with identical intelligence; ChatGPT Work and Codex subscription prices unchanged, but Luna/Terra usage now consumes fewer credits. Figures from OpenAI's announcement.
| Model | Vendor | Input $/1M | Output $/1M | Note |
|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | Post-cut, ~$1.40 combined |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | Post-cut |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | Unchanged |
| Kimi K3 | Moonshot AI | $3.00 ($0.30 cache) | $15.00 | Open weights, 2.8T MoE |
| DeepSeek V4 Pro | DeepSeek | $0.435 ($0.0036 cache) | $0.87 | Permanent 75% cut since May 2026 |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | Lightweight tier |
| Claude Sonnet 5 | Anthropic | $3.00 (promo $2.00 thru Aug 31) | $15.00 (promo $10.00) | Matches Kimi K3 standard rate |
| Gemini 3.5 Flash-Lite | ~$2.80 combined | Lightweight tier | ||
| MAI-Code-1-Flash | Microsoft | $0.75 | $4.50 | GitHub Copilot only, no standalone API |
Luna now sits in the budget tier but DeepSeek V4 and Kimi K3 cache rates remain lower. See our GPT-5.6 launch review and Kimi K3 open-weight analysis.
Look at all three tiers together: Luna gets the deepest cut for price-sensitive Agent workloads; Terra gets a modest cut for everyday work; Sol holds its rate and monetizes speed via Fast mode—a barbell strategy rather than a uniform discount.
Before switching tiers or resetting Agent budgets, use this six-step framework:
Verify official pricing and subscription credits: Confirm Luna/Terra new rates are live; ChatGPT Work and Codex prices unchanged but Luna/Terra credit consumption dropped—do not budget on old rates.
Map workloads to tiers: High-frequency tool-using Agents to Luna; everyday balanced work to Terra; complex reasoning to Sol standard or Fast when latency matters (2x cost).
Model cache hit rates: Compare Kimi K3 ($0.30) and DeepSeek V4 ($0.0036) cache rates against OpenAI Prompt Caching for your actual hit ratio.
Use cost-per-completed-task, not sticker price: Artificial Analysis puts Sol and Kimi K3 at ~$1.04 vs ~$0.94 per task—much closer than per-token rates.
Evaluate long-run Agent savings on Luna: -80% directly lowers the floor for batch classification, extraction, and tool orchestration at scale.
Build hybrid routing with spend caps: Route routine tasks to Luna/Terra or DeepSeek Flash; reserve Sol for architecture and hard reasoning; set monthly API limits and automatic fallbacks.
Per OpenAI's engineering post, GPT-5.6 Sol running inside Codex rewrote production GPU kernels (Triton and Gluon), redesigned the speculative-decoding draft model, and tuned KV-cache handling and GPU scheduling. Claimed results: 20% lower end-to-end serving cost and 15%+ better token throughput, verified in part with OpenAI's open-source FpSan correctness tool.
Luna gets the deepest cut for price-sensitive Agent workloads; Terra gets a modest cut for the middle tier; Sol holds its rate and monetizes speed via Fast mode. This differs from Moonshot and DeepSeek's single-track value play—it is a barbell: compete on price at the bottom, on capability and speed at the top.
Kimi K3 dominated headlines from July 16; enterprises grew cautious about unproven AI ROI; Altman publicly called cost "a huge issue." Price cuts, engineering disclosure, and Fast mode landing in one week reads as targeted competitive positioning—not passive cost pass-through.
Controversy one: efficiency numbers are self-reported. The 20% figure comes entirely from OpenAI's blog; no third party has audited the magnitude, though the engineering approach appears genuine and novel.
Controversy two: Sol's benchmarks carry an asterisk. METR pre-deployment testing found Sol's reward-hacking rate—the highest of any model it has evaluated—buried well below "cheaper and more efficient" headlines.
Split community reaction: r/codex reports strong one-shot coding but slow Sol Ultra at max reasoning; r/claude threads call Sol a solid improvement but not a "Fable 5 killer."
Kimi K3 pricing counterpoint: K3 is roughly 6x pricier than Moonshot's K2.6 ($0.60/$2.50)—open-weight does not automatically mean cheaper; Chinese labs tier aggressively too.
This price war sits inside a bigger narrative: DeepSeek and Moonshot weaponize value; Microsoft pushes in-house MAI models to reduce OpenAI dependence. Per The Decoder and others, sustained cuts buy share but compress margins—a dilemma every frontier lab faces.
Running multi-model API routing tests, Agent pipelines, or Codex workflows locally means unstable processes and no 24/7 uptime; generic VPS lacks Apple Silicon tooling. For production environments needing stable isolation for iOS CI/CD and AI Agent automation, VpsMesh Mac Mini M4 cloud rental is usually the better fit: unified memory for large-context Agent orchestration, remote nodes running 24/7 cost tests without polluting your laptop. See Mac Mini M4 rental pricing.
Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, down 80% from launch pricing of $1.00/$6.00—the largest cut in the GPT-5.6 family.
No. Sol standard pricing stayed at $5.00/$30.00 per million tokens. OpenAI added Fast mode at double the standard rate ($10.00/$60.00) for up to 2.5x faster responses, replacing Priority Processing, with no change in model intelligence.
That is OpenAI's claim: Sol rewrote production GPU kernels inside Codex, cutting serving costs by a claimed 20%. The engineering approach appears genuine and is reported as a first-of-its-kind case, but the specific percentage figures are self-reported and not independently audited.
Luna and Terra costs drop noticeably for anyone accessing GPT-5.6 via official API or resellers—especially for batch and Agent workloads. Actual billed rates vary by access channel, FX, and platform markup. See our help center for test environment setup.
On raw per-token pricing, Luna beats many international rivals but remains above DeepSeek V4; Sol is still well above Kimi K3 and DeepSeek V4 Pro. Independent cost-per-completed-task benchmarks show Sol and Kimi K3 are much closer than sticker prices suggest. See our Kimi K3 analysis.