\n\n\n\n Half a Trillion Parameters, Twenty-Three Billion Doing the Work - Agent 101 \n

Half a Trillion Parameters, Twenty-Three Billion Doing the Work

📖 5 min read•826 words•Updated Oct 5, 2026

Beam is the biggest model in its comparison group and also the smallest. Both things are true, and the gap between them is the whole story.

Reflection AI announced Beam on October 5, 2026, calling it the company’s first open-weight model. The headline number is 501 billion parameters. The number that actually matters is 23 billion. If you’ve been nodding along politely whenever someone mentions “mixture of experts” at a meeting, this is a good moment to figure out what that phrase means, because Beam is built almost entirely around it.

What a mixture of experts actually does

Picture a large hospital. It has hundreds of specialists on staff, but when you walk in with a sprained ankle, you don’t get examined by all of them. You see two or three people. The rest are on payroll, available, and completely irrelevant to your ankle.

That’s a sparse Mixture-of-Experts model. Beam holds 501 billion parameters in total, but only around 23 billion of them switch on for any given request. A routing system picks which “experts” inside the model are relevant and ignores the rest.

Why does anyone care? Because the cost of running a model scales with how much of it you actually wake up, not how much of it exists. Storage is cheap-ish. Computation is not. A model with a huge total size and a small active size gives you deep specialized knowledge without paying to run the whole thing on every question.

Reflection says Beam competes with larger open models like GLM 5.2 while using three to four times less inference compute on reasoning benchmarks. For a non-technical reader, that’s the sentence to remember. Comparable results, a fraction of the compute to get there.

Why it’s described as both big and small

In the comparison set Reflection published, Beam has the fewest total parameters of the group. Its 23 billion active parameters also sit below GLM-5.2 and Nemotron 3 Ultra. So the half-trillion figure that sounds enormous in isolation is, in context, the modest option. Model size numbers have inflated to the point where 501 billion reads as restrained, which tells you something about where this space has gone.

Built for agents, not conversation

Reflection positions Beam for coding, reasoning, and agentic workloads. That last term is the one worth unpacking on a site like this.

A chatbot answers your question and stops. An agent is given a goal and then takes a series of steps on its own: reading files, running commands, checking whether the result worked, trying again when it didn’t. Agents don’t make one model call. They make dozens, sometimes hundreds, in a single task.

That changes the economics completely. If a model costs a little more per call, a chat product barely notices. An agent running a 200-step task notices enormously. This is why the active-parameter count is the number builders will fixate on. A design that uses three to four times less compute per step is the difference between an agent you can afford to run and one you can only demo.

On capability, Reflection reports 80.9% on SWE-bench Verified. SWE-bench is a test built from real software bugs pulled from real open-source projects. The model gets the codebase and the bug report, and has to produce a fix that passes the project’s own tests. The “Verified” version is a human-checked subset, which makes it one of the more trustworthy coding measures available. These are scores as published by Reflection, not independently reproduced.

The licensing is the quiet part

Beam’s weights are planned for release under Apache 2.0 and MIT licenses. Those are two of the most permissive licenses in software. In plain terms: you can download the model, use it commercially, modify it, and build products on top of it without asking permission or sharing revenue.

Many models marketed as “open” come with restrictions on commercial use, user-count caps, or clauses blocking you from training competing models. Apache 2.0 and MIT have none of that. For a company releasing a model at this scale and claiming these benchmark results, choosing the permissive option is a real decision with real consequences.

What you can’t do yet

You can’t download it. Reflection says it still needs to finish final red-teaming, which is adversarial safety testing where people actively try to make the model misbehave. The plan is to publish the weights, a technical report, a model card, and developer tools across the rest of October 2026. For now, access runs through an early access waitlist.

So the honest status: announced, benchmarked by its maker, not yet verifiable by anyone else. If you’re evaluating Beam for actual work, the useful move is to wait for the weights and the technical report, then check whether independent testing matches the published numbers.

What’s already clear is the design philosophy. Reflection didn’t chase the largest possible model. It built one where most of the parameters stay asleep most of the time, and bet that efficiency per step is what agent builders will pay for. That bet looks well-aimed.

🕒 Published:

🎓
Written by Jake Chen

AI educator passionate about making complex agent technology accessible. Created online courses reaching 10,000+ students.

Learn more →
Browse Topics: Beginner Guides | Explainers | Guides | Opinion | Safety & Ethics
Scroll to Top