\n\n\n\n When Memory Gets Scarce, AI Gets Expensive - Agent 101 \n

When Memory Gets Scarce, AI Gets Expensive

📖 5 min read•828 words•Updated Sep 10, 2026

AI chips just got pricier in China.

According to a Reuters exclusive dated September 10, Chinese AI chipmakers including Huawei Technologies and Cambricon have sharply raised prices on their AI processors. Three people familiar with the matter described increases of roughly 20 to 50 percent compared with quotes given just a couple of months earlier. Huawei’s Ascend 950DT accelerator card, which packages an AI processor together with memory and other components, now carries an indicated price above 250,000 yuan, or about $37,255.

The culprit is a shortage of something most people have never heard of: high-bandwidth memory, usually shortened to HBM.

What high-bandwidth memory actually does

If you’ve ever wondered why AI hardware costs so much more than a laptop, memory is a bigger part of the answer than most explanations let on.

Think of an AI chip as a very fast chef. The chef can chop, sear, and plate at superhuman speed. But the chef can only work as fast as ingredients arrive from the pantry. If someone is walking one carrot at a time down a long hallway, all that knife skill is wasted.

High-bandwidth memory is the pantry moved into the kitchen, with a wide conveyor belt feeding the counter. It sits physically close to the processor and shuttles enormous amounts of data back and forth. For AI work, that matters more than it does for almost any other kind of computing, because large models are, at heart, giant piles of numbers that have to be read repeatedly and quickly.

Which means an AI accelerator isn’t really one product. It’s a processor plus its memory, sold as a unit. When the memory gets expensive or hard to find, the whole card gets expensive. That’s exactly what Huawei’s price on the Ascend 950DT reflects.

Why Huawei built its own

Huawei has said the Ascend 950 series will use two of its own HBM technologies: HiBL 1.0 in the Ascend 950PR and HiZQ 2.0 in the Ascend 950DT. The company hasn’t disclosed the technical details of either.

Proprietary memory is an unusual choice. It’s a significant undertaking to develop, and it tells you how central memory has become to whoever wants to build AI hardware at scale. It also hasn’t insulated anyone from the price pressure, since the shortage is being felt across current and next-generation parts alike.

What this means for AI agents

Readers here mostly care about AI agents: the systems that book things, answer things, research things, and take actions on your behalf. Chip pricing feels several floors below that. It isn’t.

Every agent action has two cost centers behind it. There’s training, which is the expensive one-time-ish work of building the model. And there’s inference, which is what happens each and every time your agent thinks. Reuters reports these price increases are affecting the cost of both AI model development and inference.

Inference is the one to watch, because agents are unusually hungry there. A chatbot answers your question once. An agent might reason through a plan, check a result, revise, call a tool, and reason again. One user request can become dozens of model calls. That multiplication is what makes agents useful, and it’s also what makes their economics sensitive to hardware costs in a way that simpler AI products aren’t.

So when the underlying silicon gets 20 to 50 percent more expensive, the pressure lands hardest on exactly the kind of product that runs a model many times per task.

The likely knock-on effects

Nothing here is predicted in the reporting, so treat this as my read rather than fact. But the pressure points are fairly predictable when compute gets pricier:

  • More use of smaller, cheaper models for routine steps, with the expensive model reserved for the hard parts
  • Tighter limits on how many reasoning steps an agent is allowed to take before it has to give an answer
  • More caching, so identical requests don’t get recomputed from scratch
  • Slower rollouts of features that would work but can’t yet be run affordably at scale

Most of that is invisible to the person using the product. You don’t see the routing decision. You just notice that the agent feels a bit more clipped, or that a feature stayed in beta longer than expected.

The unglamorous part of the story

AI coverage tends to focus on the models: what they can write, what they can reason through, what benchmark they cleared this month. The physical supply chain gets far less attention, right up until it becomes the constraint.

Memory is that constraint at the moment, at least for chipmakers in China. It’s not the sort of thing that makes for a good demo. It’s a component, in short supply, priced accordingly, and the cost is working its way up through model development and into inference.

Worth remembering the next time an agent product’s pricing shifts, or a promised capability arrives later than announced. The explanation may have less to do with the model and more to do with the conveyor belt feeding it.

🕒 Published:

🎓
Written by Jake Chen

AI educator passionate about making complex agent technology accessible. Created online courses reaching 10,000+ students.

Learn more →
Browse Topics: Beginner Guides | Explainers | Guides | Opinion | Safety & Ethics
Scroll to Top