OpenAI cuts GPT-5.6 Luna by 80% — from $1 to $0.20 per million input tokens. Three weeks after launching it. The fastest major price cut in the history of commercial AI.
That's the headline.
Here's the context that makes it meaningful for small businesses: this isn't a one-off promotional discount.
It's the signal that the AI pricing war between US labs and Chinese competitors has reached a tipping point and the winner, by a considerable margin, is every small business that uses AI for anything.
In our recent post (Issue #23), we wrote about how chinese AI models are far cheaper than their western counterparts and then this update happens.
🔧 What Actually Happened and Why It Matters
On July 30, 2026, OpenAI announced two pricing changes across the GPT-5.6 family:
GPT-5.6 Luna: $1 → $0.20 per million input tokens. $6 → $1.20 per million output tokens. An 80% reduction, three weeks after launch.
GPT-5.6 Terra: $2.50 → $2.00 per million input tokens. $15 → $12 per million output tokens. A 20% reduction.
GPT-5.6 Sol: Unchanged at $5/$30.
OpenAI attributed the cuts to efficiency improvements — GPT-5.6 Sol had been used to optimise their own production infrastructure, cutting serving costs by 20% and improving token generation efficiency by over 15%.
The AI was used to make the AI cheaper to run. That's the loop now.
But the real driver, as CNBC reported independently, is competition.
Chinese models like DeepSeek V4 Pro at $0.435/$0.87, Kimi K3, and others had captured 46% of US enterprise token usage on OpenRouter. OpenAI's Luna cut, at $0.20 input, now undercuts DeepSeek on input costs, though DeepSeek remains cheaper on output tokens.
The result is a market where frontier-quality AI text processing now costs a fraction of what it did 12 months ago:
Model | Input (per 1M tokens) | vs. 12 months ago |
|---|---|---|
GPT-5.6 Luna | $0.20 | ~85% cheaper |
GPT-5.6 Terra | $2.00 | ~60% cheaper |
DeepSeek V4 Flash | $0.14 | Never existed |
Claude Sonnet 5 | $2.00 | ~40% cheaper (intro) |
Gemini 3.6 Flash | Competitive | Significantly cheaper than 2025 |
The era of "AI is expensive" is over.
What replaced it is an era where the cost of running AI at meaningful business volume has dropped below the cost of one employee hour.
🧪 Real Business Example
A small digital marketing agency was spending $340/month on API costs across their content production workflow using Claude Sonnet 4.6 for first drafts, GPT-5.5 for editing passes, and a separate model for SEO analysis.
Their monthly output: approximately 120 pieces of content for clients.
After the Luna price cut and reassigning tasks by model tier, their revised stack: Luna at $0.20 for first-draft generation (high volume, lower stakes), Terra at $2 for editing and refinement passes (moderate volume, higher quality bar), Fable 5 for complex strategic documents (low volume, highest stakes).
Same output volume. Monthly API spend: $89 — a 74% reduction.
They passed none of it on to clients. It went straight to margin.
Their observation after one month: the quality difference between Luna's $0.20/million drafts and their previous $1/million setup was not detectable in client-facing work for the content categories they produce. The pricing changed. The output quality, for their use case, didn't.
📋 Step-by-Step: Audit Your AI Costs and Reassign by Task
List every AI tool you pay for — subscriptions and API costs separately. Include anything with AI baked in: your writing tool, your customer service chatbot, your content scheduler.
Categorise your AI tasks by stakes — high-stakes (client proposals, strategic documents, legal or financial content), medium-stakes (marketing copy, email drafts, analysis), low-stakes (first drafts, summarisation, classification, repetitive generation).
Match model tier to task stakes — the principle: use the cheapest model that produces acceptable output for each task type.
Low-stakes, high-volume: Luna ($0.20) or DeepSeek V4 Flash ($0.14)
Medium-stakes: Terra ($2) or Claude Sonnet 5 ($2 intro)
High-stakes: Fable 5, Sol, or Opus 4.8 — reserve for tasks that genuinely need frontier reasoning
Test the cheaper model on your actual tasks — don't assume quality difference without testing. Run the same prompt through Luna and your current model. Compare outputs. For many content tasks, the gap is smaller than the price gap.
Rebuild your Zapier/Make automations to use Luna or Flash — if you have automations that route through GPT-4 or an older model by default, update the model string to Luna or an equivalent. The efficiency gain on high-volume automations is immediate.
Calculate your new monthly cost — multiply your typical monthly token volume by the new rates. If you're on a flat $20/month subscription rather than API billing, the savings are indirect — you get more tokens for the same budget, meaning higher usage limits before hitting caps.
Redirect the savings — be deliberate about where the efficiency gain goes. More content volume? Same volume with higher margin? Reinvestment in the one or two frontier-model tasks that genuinely need Sol or Fable 5?
❓ The Dumb Question
"If AI is getting so cheap, why does my $20/month ChatGPT subscription cost the same?"
Because subscription pricing and API pricing are different products.
Your $20/month Plus subscription gives you a usage-limited interface with no per-token billing.
OpenAI's API pricing, the thing that dropped 80% is for developers and businesses building on top of ChatGPT programmatically at volume.
The two are related but separate. The subscription price didn't change.
What changed is: within your subscription budget, the same token limits now buy more efficient processing which means slightly higher effective usage limits before you hit caps.
More practically: if you're a small business owner using ChatGPT via the website or app, the 80% cut matters less to you directly than it does to a business running AI through the API at scale.
For most small business owners, the more immediately relevant implication is that the AI tools you'll adopt over the next 12 months: chatbots, automation platforms, custom AI features in your existing software might be significantly cheaper to power than they would have been 12 months ago.
💰 The Current AI Model Price Map (August 2026)
Model | Input (per 1M tokens) | Output (per 1M tokens) | Best For |
|---|---|---|---|
DeepSeek V4 Flash | $0.14 | $0.28 | Highest volume, cost-critical |
GPT-5.6 Luna | $0.20 | $1.20 | High volume, OpenAI ecosystem |
GPT-5.6 Terra | $2.00 | $12.00 | Balanced, production use |
Claude Sonnet 5 | $2.00 | $10.00 | Writing quality, intro pricing until Aug 31 |
DeepSeek V4 Pro | $0.435 | $0.87 | Reasoning at low cost |
Claude Fable 5 | $10.00 | $50.00 | Frontier reasoning, Max/Team plans |
GPT-5.6 Sol | $5.00 | $30.00 | Complex reasoning, agentic work |
Note: Claude Sonnet 5's introductory pricing ends August 31, 2026. If you're building on Sonnet 5 via API, lock in now before standard rates apply.
⚡ The Practical Play
This week: if you use any AI tool through an API or pay-per-use model, pull up your last invoice and calculate your cost per task.
Then check whether GPT-5.6 Luna at $0.20 or DeepSeek V4 Flash at $0.14 can handle that same task at acceptable quality.
A 20-minute test against your actual prompts and a cost comparison is all it takes to know whether a model switch pays off.
The savings are real. The test is free.
📰 News That Matters
The numbers behind the ChatGPT 1 billion user announcement are worth sitting with. OpenAI says users send 50% more messages per day after six months of use and use ChatGPT for twice as many types of tasks.
That's the adoption curve in real-time: people start using AI for one thing, find it works, and expand to more.
Agentic work through Codex now accounts for 99.8% of weekly output tokens across OpenAI itself meaning almost all of OpenAI's own AI usage is autonomous task completion, not Q&A.
That's the direction the whole industry is heading.
The price cuts accelerate it: when frontier AI costs $0.20/million tokens, the case for building it into every repetitive business workflow becomes almost impossible to argue against.
🚫 Skip This
Chasing the absolute cheapest model for every single task.
The efficiency gain from tiered model selection is real but taking it too far produces a different problem: time spent managing model selection, evaluating outputs from cheaper models, and catching errors that a better model would have avoided.
The goal is appropriate model for appropriate task, not always cheapest model everywhere.
Luna and Flash are excellent for high-volume, pattern-consistent tasks. They're not the right choice for client proposals, complex analysis, or anything where a single error has material consequences.
The price war is a gift. Use it strategically, not indiscriminately.
Until next issue, Kris
The Layman's AI — The only AI updates your business actually needs.
