If you run AI agents, the most important thing OpenAI shipped this month is not a smarter model — it’s a cheaper one.
On September 22, 2026, OpenAI released GPT-6 Sol and GPT-6 Luna, two models that slot in as direct replacements for the GPT-5.6 lineup at half the price. They’re the cheaper, faster siblings to GPT-6 Astra, the flagship that arrived earlier in the month. Astra got the headlines. Sol and Luna are the ones that will quietly change what you can actually afford to build.
What the new prices actually are
Here’s the pricing, and it’s the whole story:
- GPT-6 Sol — $2 per million input tokens, $10 per million output tokens. That’s down 50% from GPT-5.6 Sol’s promotional prices.
- GPT-6 Luna — $0.10 per million input tokens, $0.50 per million output tokens.
If “tokens” makes your eyes glaze over, think of them as the AI equivalent of words. Roughly speaking, a million tokens is a small library shelf of text. So Luna reads a shelf for a dime and writes back for fifty cents. Sol costs more because it’s the stronger of the two, but $2 to read a million tokens is the kind of number that used to be reserved for the weakest models available.
Why agent builders should care more than chatbot users
This is the part that gets missed. If you’re a person typing into a chat window, price cuts are nice but mostly invisible. You send a few hundred words, you get a few hundred back, and your bill is rounding error.
Agents don’t work that way. An agent is a model that runs in a loop — it reads a task, thinks, calls a tool, reads the result, thinks again, and repeats until it’s done. Every single pass through that loop re-sends context. A single “go research this and write me a summary” request can involve twenty model calls and chew through vastly more tokens than a human conversation ever would.
Which means agent economics are absurdly sensitive to per-token price. Cut the price in half and you haven’t made your agent 50% cheaper to run — you’ve made a whole category of agents that were too expensive to justify suddenly viable. The background task that checks your inbox every fifteen minutes. The agent that reviews every support ticket instead of a sample. The one that tries three approaches and picks the best result. Those were all calculator-out-loud decisions before. At these prices, a lot of them stop being decisions at all.
A million tokens of memory, on both models
Both Sol and Luna carry a 1.05M-token context window, of which 922K can be input and 128K is the maximum output. Context window is just how much the model can hold in its head at once. Nearly a million tokens of input is a genuinely large working memory — enough for long document sets, sprawling conversation histories, or the accumulated notes an agent builds up over a long task.
What stands out is that the cheap model gets the same window as its more capable sibling. Historically, budget models came with cramped context, which made them useless for exactly the long-running agent work where you most wanted to save money. Not the case here.
One small difference worth knowing: their knowledge cutoffs aren’t identical. Sol’s training knowledge runs to April 20, 2026, and Luna’s to May 18, 2026. That’s the date past which each model simply doesn’t know what happened. If your agent needs current information, that gap is what search tools and retrieval exist to fill — regardless of which model you pick.
Sol or Luna, practically speaking
The split between them is the useful design decision. Sol is the stronger, pricier option. Luna is roughly twenty times cheaper on input. In agent terms, that maps pretty cleanly onto a two-tier setup: let Luna handle the high-volume grunt work — classifying, filtering, summarizing, deciding whether something even deserves attention — and escalate to Sol when the task needs real reasoning. Astra sits above both for the hardest problems.
That layering isn’t new as an idea. What’s new is that the cheap tier is cheap enough, and holds enough context, that you can be genuinely wasteful with it. Run it on everything. Run it twice.
The competitive angle
Coverage of the launch framed it as undercutting Claude on price, and that read seems fair. Model pricing has become a market where nobody holds a lead for long, and a 50% cut across a lineup is a fairly direct signal about where OpenAI expects the fight to be.
For anyone building with these tools, that competition is unambiguously good news. The models keep getting better and the meter keeps getting slower. Sol and Luna aren’t exciting in the way a flagship release is exciting. They’re exciting in the way a utility bill dropping by half is exciting — quietly, and in ways you’ll feel every month.
🕒 Published: