\n\n\n\n Sonnet 5 and the Discount That Refused to Expire - Agent 101 \n

Sonnet 5 and the Discount That Refused to Expire

📖 5 min read•819 words•Updated Sep 28, 2026

Anthropic told everyone its new pricing on Claude Sonnet 5 would last until August 31, 2026, after which the cost would climb to $3 per million input tokens and $15 per million output tokens. On August 10, three weeks early, the company made the cheaper rate permanent instead.

Those two facts sitting next to each other tell you more about where AI agents are heading than any launch video. A company that plans to raise prices does not usually lock in the lower one ahead of schedule. So let’s unpack what actually happened, and what it means if you’re someone who uses these tools rather than builds them.

What Anthropic shipped

Claude Sonnet 5 arrived on June 30, 2026. Anthropic described it as its most agentic Sonnet model to date, built to plan multi-step tasks and operate tools with less human supervision. It carries a 1M-token context window. On July 1, it became the default model for all Free and Pro users, replacing Sonnet 4.6.

That default switch is the part most people will feel without noticing. If you opened Claude on July 1 and typed a question, you were talking to a different model than the day before. No setting to change, no upgrade to buy.

Translating the jargon

A few terms worth having in plain language:

  • Tokens are chunks of text, roughly three-quarters of a word each. Pricing is quoted per million of them because that’s the unit the model reads and writes in. Input tokens are what you send; output tokens are what it sends back. Output costs five times more here, which is why long-winded answers are expensive and long documents are relatively cheap to read.
  • Context window is how much the model can hold in working memory at once. A million tokens is a lot of material, several long books’ worth, which matters when you want an agent to keep track of a project instead of a single question.
  • Agentic means the model is designed to do things in sequence rather than just answer. Break a goal into steps, call a tool, look at the result, decide what’s next. The phrase “reduced human supervision” in Anthropic’s own description is the operative bit.

Why cheap and agentic go together

These two design choices are not separate features that happened to ship on the same day. They depend on each other.

A model that answers one question uses a predictable amount of tokens. A model that plans a task, tries something, reads the outcome, and adjusts uses far more, because every loop feeds the previous result back in as new input. The same job that costs a few cents as a single exchange can cost many times that as an agent working through it step by step.

Which means price is not a marketing detail for agent work. It’s a ceiling on what’s practical. At $2 in and $10 out, a mid-tier model doing twenty rounds of tool calls stays in territory most teams can justify. Push that to $3 and $15 and some of those workflows stop making sense on a spreadsheet.

Anthropic also priced Sonnet 5 below its flagship Opus model through August 31. Read those two moves together and the strategy comes into focus: put the agentic model at a price point where people will actually let it run, rather than reserving that behavior for the expensive tier.

What this means if you’re not a developer

You might reasonably ask why token prices matter to you if you pay a flat Pro subscription. Fair question. The answer is that the tools you’ll be offered over the next year are shaped by what they cost to run.

Every AI feature bolted onto a product you already use, the assistant in your email, the helper in your project tracker, the thing that drafts your meeting notes, is somebody paying per token behind the scenes. When those costs drop, the calculation changes. Features that were too expensive to offer become viable. Agents that were limited to three steps get room for thirty.

The flip side deserves attention too. “Reduced human supervision” is a selling point and a risk in the same breath. A model that takes twenty actions on your behalf is a model that can make twenty mistakes before you look. Cheap autonomy makes it easier to hand over tasks you haven’t thought carefully about handing over.

My practical advice hasn’t changed much: start agents on work where you can check the output quickly and where a wrong answer is annoying rather than costly. Scheduling, drafting, summarizing, sorting. Save the high-stakes stuff for when you’ve watched the thing work for a while and know its habits.

Anthropic keeping its lower price early is a small signal with a clear direction. The company that told us its agentic model was worth $3 per million tokens decided $2 was the number that mattered more. That’s not generosity. That’s a bet on volume, on agents running long and often, and on a lot more of us letting them.

🕒 Published:

🎓
Written by Jake Chen

AI educator passionate about making complex agent technology accessible. Created online courses reaching 10,000+ students.

Learn more →
Browse Topics: Beginner Guides | Explainers | Guides | Opinion | Safety & Ethics
Scroll to Top