\n\n\n\n Twenty-Five Million Dollars for Very Educated Guesses - Agent 101 \n

Twenty-Five Million Dollars for Very Educated Guesses

📖 5 min read•821 words•Updated Sep 19, 2026

The forecasters got out-forecast.

On September 18, a London startup called Mantic announced it had raised $25 million in seed funding. The reason investors opened their wallets is unusual, and honestly kind of fun: Mantic’s AI system entered a public forecasting tournament, the summer 2026 Metaculus Cup, and beat every human who showed up. It finished second overall. The only competitor ahead of it was another bot.

If you’re not steeped in this world, that sentence might land softly. Let me explain why it’s a bigger deal than it sounds.

What forecasting tournaments actually are

A forecasting tournament isn’t a trivia quiz. Nobody is asking who won the 1998 World Cup. Participants are asked about things that haven’t happened yet, and instead of answering yes or no, they answer with a probability. Will a particular election go a certain way? Will an economic indicator cross a threshold by December? Will a cultural event play out the way people expect?

You say something like “68% chance.” Then reality happens, and your score depends on how well-calibrated you were. Confidently wrong answers hurt a lot. Vague hedging at 50% on everything gets you nowhere. To score well over dozens of questions, you have to actually be right about the world, repeatedly, with the right amount of certainty attached.

The people who compete in these things are serious. Superforecasters, as they’re often called, have a genuine track record of beating analysts and pundits. They read widely, update their views when news breaks, and treat their own confidence as something to be measured rather than felt. Mantic’s system, per the Metaculus Cup results, assigned probabilities to political, economic and cultural questions and came out ahead of all of them.

Why this is an AI agent story, not just an AI story

This is the part I find most interesting for readers of this site. A chatbot answers questions from what it already knows. A forecasting system has to do something structurally different. To predict whether something will happen next month, it needs to go find current information, weigh sources against each other, notice what changed since yesterday, and convert all that into a single number it’s willing to be graded on.

That loop of gathering, weighing, and committing is roughly what people mean when they say “AI agent.” The output isn’t prose. It’s a decision with a confidence level attached, and it can be checked against reality later. That last part matters more than anything else here. Most AI claims are hard to verify because “good writing” and “helpful answer” are subjective. A probability on a future event is not subjective. Either the thing happened or it didn’t, and the scoreboard is public.

Mantic’s win is a rare case of an AI system being graded in the open, against motivated humans, on a task where you can’t fake it.

Follow the money

According to the reporting, Mantic drew interest from global companies and government agencies, with hedge funds and trading firms showing particular enthusiasm. That last group tells you something.

Trading firms don’t buy forecasts because they’re intellectually charming. They buy them because a small, consistent edge in estimating probabilities is worth real money. If a system can tell you that the market is pricing an event at 40% when the honest number is 30%, that gap is the entire business. Hedge funds noticing a forecasting startup is the clearest signal available that people think the output is useful rather than merely impressive.

Government agencies are interested for a different reason. Policy planning is forecasting, whether or not anyone calls it that. So is supply chain planning, insurance pricing, and deciding whether to open an office in a particular country next year.

The healthy skepticism section

A few things worth holding onto before we get carried away.

  • One tournament is one tournament. A strong showing in a single competition is evidence, not proof. Calibration is measured over long stretches precisely because short runs can flatter you.
  • Mantic finished second, behind another bot. The headline is that AI beat the best humans in that field, not that Mantic is the best forecaster in existence.
  • Tournament questions are written to be resolvable. Real business and policy questions are often messy, poorly defined, and lack a clean resolution date. Performance in a scored contest doesn’t automatically transfer.
  • $25 million is seed funding. It’s a bet on what the team can build, not a verdict on what already exists.

What I’d watch for

The thing I’ll be tracking isn’t Mantic’s next funding round. It’s whether these systems stay calibrated once the questions get harder and the stakes get real. An agent that knows how confident it should be is genuinely more useful than one that sounds confident all the time, and it’s a quality almost nobody measures in consumer AI products.

If forecasting tournaments become the standard way we grade AI agents, that would be a solid outcome for everyone. Public scoreboards are harder to spin than demos.

🕒 Published:

🎓
Written by Jake Chen

AI educator passionate about making complex agent technology accessible. Created online courses reaching 10,000+ students.

Learn more →
Browse Topics: Beginner Guides | Explainers | Guides | Opinion | Safety & Ethics
Scroll to Top