Microsoft taking down EvilTokens is not really a story about Microsoft winning. It is a story about how cheap it has become to attack people at scale, and that part did not get fixed last week.
Here is what we know. Microsoft disrupted a platform called EvilTokens, an AI-assisted service that was used to compromise 12,000 accounts. The platform offered what reports describe as a streamlined, end-to-end service for cybercriminals, automating various stages of an attack. Microsoft’s action was meant to dismantle this specific operation and send a signal to anyone running something similar. The available reporting does not give a precise disruption date or much on what happened afterwards.
That is a short list of facts. But if you read it slowly, the uncomfortable detail is sitting right there in the middle: end-to-end service.
What “end-to-end service” actually means
I spend most of my time explaining AI agents to people who do not write code, so let me translate that phrase, because it is doing a lot of work.
An AI agent is software that takes a goal and then carries out the steps toward it on its own. You give it an objective, it figures out the sequence, and it keeps going without someone clicking a button at each stage. That is the useful version: an agent that sorts your inbox, books a meeting, chases down a refund.
Now apply the same structure to an attack. A traditional cyberattack has stages. Find targets. Write something convincing enough to fool them. Deliver it. Collect whatever comes back. Do something with the access you gained. Historically each of those stages took a person with a specific skill, and that was a real bottleneck. Writing a convincing phishing message in fluent English is a skill. Knowing which 5,000 email addresses are worth trying is a skill.
“Automating various stages” means those bottlenecks got widened. “End-to-end service” means someone packaged the whole chain up and rented it out. The customer does not need the skills anymore. They need a budget.
The number that matters is not 12,000
Twelve thousand accounts is a lot of people having a bad week. Still, it is not an unprecedented number in security terms. Breaches have hit far larger figures.
What is different is the ratio. The question I would want answered is not how many accounts fell, but how few people it took to knock them over. That is the thing agent-style automation changes. It decouples the size of an attack from the size of the team running it. One operator with a credit card and a service subscription can now do work that used to need a crew.
When that ratio shifts, the economics shift too. Attacks that were not worth the effort become worth the effort. Targets that were too small to bother with become profitable. If you have ever thought “why would anyone come after me, I’m nobody interesting,” that reasoning was always a bit shaky, and automation at this level quietly retires it.
Why the takedown is good and also not enough
I want to be fair to Microsoft here. Disrupting an operation like this takes real work, and the deterrent value is genuine. If you are considering building the next EvilTokens, you now know a company with Microsoft’s resources may come for your infrastructure. That matters.
But takedowns treat the supplier, not the demand. The demand is sitting there, unchanged, with money in hand. And the discussion in the Ars OpenForum thread about this story raised a point I keep coming back to: if an operation runs on open-weight models, the ones anybody can download and run on rented hardware, there is no account to suspend and no vendor to notify. The capability does not live on anyone’s server.
That is not an argument against open models. It is an observation about where the use sits. Removing one platform from the internet is a legitimate win. It is not the same as removing the capability from the world.
What this means if you are not in security
You do not need a threat model. You need to update one assumption, and it is this: the sloppy phishing email with the weird grammar and the obviously fake sender name was never the ceiling of what attackers could do. It was a symptom of the work being manual and the attacker being in a hurry.
Take that constraint away and the messages get more fluent, more specific, and better timed. The old advice to watch for typos was never great, and it is now closer to useless.
What still works is boring and structural. Multi-factor authentication, because stolen passwords alone stop being enough. A password manager, so one compromised account does not cascade. And a habit of verifying unexpected requests through a different channel than the one they arrived in.
EvilTokens is gone. The reason it worked is not. That is the part worth sitting with.
🕒 Published: