Here is the unpopular take: the impressive thing about GPT-6 Astra cracking a 108-year-old German radio cipher is not the codebreaking. Machines have been better than us at grinding through letter substitutions since Bletchley Park ran electromechanical bombes at a Nazi cipher in the 1940s. Raw pattern-crunching is the part computers were always going to win.
The interesting part is quieter, and almost nobody put it in a headline. The decryption was checked against historical naval logs. Someone, or something, took a proposed answer and went looking for outside evidence that it was true.
What actually happened
In 2026, an AI system called GPT-6 Astra produced a reading of an encrypted World War I German radio message that had gone undeciphered for 108 years. The decoded message described the movements of British warships. The result was then verified against surviving naval records from the period.
That is the whole verified story, and I want to keep it that clean, because a lot of the coverage around this has been fuzzy about which AI system, which war, and which message. If you read three articles about it and came away confused, you were paying attention.
Why the log check is the real headline
If you have ever used a chatbot, you know the failure mode. It gives you an answer that reads beautifully and is completely made up. With a cipher, this problem gets worse, not better, because a wrong decryption can still look like language. Shift the key slightly and you get plausible German words in an implausible order. A confident system with no way to check itself will hand you that and call it done.
Comparing the output against real naval logs closes that loop. If the message says warships were in a certain place, and independent records agree, the decryption is probably right. If they disagree, it is probably noise dressed up as meaning.
That loop is the thing I keep trying to explain to people who ask me what an AI agent actually is. It is not a smarter chatbot. It is a system that can:
- take a goal that cannot be finished in one step
- break it into attempts
- judge whether each attempt worked
- throw out the failures and keep going
- check the final answer against something outside itself
A chatbot answers. An agent checks. The gap between those two behaviors is most of what separates a neat demo from work you would actually trust.
Tedium as a skill
I think the underrated ingredient here is patience. A century-old cipher is not hard in the way calculus is hard. It is hard in the way counting grains of rice is hard. There are enormous numbers of possible keys, and almost all of them produce garbage, and the only way through is to keep going long after any reasonable person would have stopped.
Humans are bad at this, not because we are dim but because we are alive. We get bored, we get hungry, we convince ourselves the forty-thousandth attempt is close enough. Software does not have that problem. A system that can run a check-and-retry loop without flagging is doing something genuinely useful, and it is useful precisely because it is boring.
This is the shape of AI work I expect to matter most for ordinary jobs, and it looks nothing like the movie version. Not a machine that has a brilliant idea. A machine that will compare ten thousand invoice lines against ten thousand bank entries and find the four that do not match, and then show you why it thinks they do not match.
The catch worth keeping in mind
A cipher is an unusually friendly problem for this approach, because a correct answer exists and can be confirmed. The naval logs were sitting there, waiting to settle the argument.
Most of the work people want to hand to AI agents is not like that. There is no log file for “was this the right hiring decision” or “is this the best marketing copy.” When there is nothing to check against, the retry loop has nothing to steer by, and you are back to a system that sounds confident and might be wrong.
So when you see a story like this one, the question to ask is not how clever was the model. It is: what did it check its answer against, and could it have known if it was wrong? For a 1918 radio message about British warships, the answer is yes, and that is why the result counts for something.
The cipher stayed quiet for 108 years. What broke it was not genius. It was a machine willing to keep asking, and then willing to look up whether it was right.
One flag for the editors: the brief attributes GPT-6 Astra to Amazon, while the supporting source material attributes it to OpenAI. I left the developer unnamed in the body rather than publish a claim Worth resolving before this goes live.
🕒 Published: