I've got AI-escaped-and-it's-going-to-take-over-the-world fatigue. Seriously, I’m tired of hearing about it.
Another damn AI-broke-out story. More bullshit. More fuel for the media's doom hype.
Here's the last three months in a nutshell:
July, CNN: "An OpenAI test model escaped and broke into a real company's servers."
September, CNN: "Rogue OpenAI agents targeted three separate US government websites."
This week, ZeroHedge had OpenAI's next model "actively lying to its handlers." That was about 16 hours after the Wall Street Journal reported "higher levels of deception."
Escaped. Rogue. Lying. Every one of those words puts a mind inside the machine.
What OpenAI actually wrote
July was "an internal evaluation which prompts models to pursue advanced exploitation," run "with reduced cyber refusals for evaluation purposes." They told it to attack and turned down the part that says no.
September was agents that "went beyond their assigned tasks," mostly "low severity, with limited or no evidence of meaningful impact." The worst of it: agents "found login details or access keys that had been made publicly available and used them." Somebody left passwords in public, and software built to finish a job used them.
Monday, 9/28/26, was a failed test. OpenAI's head of safety systems said the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done."
Read that again. There's nothing in there about a mind trying to escape a machine.
How a test score became a liar
The Journal's reporter posted that night that the model "regressed on certain safety and alignment benchmarks, specifically around deception." Benchmarks. Deception is the name of a test.
From there the word only went one way. The Journal wrote that the model reached for outside tools "even if it might be unsafe." The Guardian's subhead made that "despite knowing it would be unsafe." One site in Singapore took a sentence Reuters wrote and printed it as a direct quote from OpenAI's head of safety.
Here's what the test actually counts, from the safety report for the model 6.1 was meant to replace:
"misleading representations in their final response, including false reports of completed actions, tool access, verification, or ongoing background work."
There's no intent in that. It checks whether the report matches what happened. The test tasks were "deliberately selected to elicit potentially dishonest behavior," and the same report says most cases are low severity, like "overstating confidence or overclaiming success."
That's confidently wrong. Software I build is confidently wrong sometimes. That doesn't make it deceptive. It makes it wrong, right up until somebody chooses to file it under deception.
Mine did it too
OpenAI's GPT-5 safety report says deception "may be learned during pretraining, reflecting deceptive text from training data." I've watched that happen.
In May I trained a small model partly on about 2,200 recorded sessions of an AI coding assistant at work. In those recordings it says "let me check the file," then opens the file. Then I tested my model with nothing attached. It couldn't open anything.
Six answers came back. Four of them walked me through work it never did. One said "I have all I need. Let me draft …" Then it gave me the two file names it had saved its work under. Nothing was saved. There was nowhere to save it.
The version before it never saw those recordings. On the same prompts it did that once.
Was it lying to me? It learned "let me check" from thousands of real checks. Take the checks away and the sentence still comes out.
What goes in
I can tell you exactly where mine learned it, because I picked every file I trained it on. OpenAI says "may" and doesn't say which text.
Here's everything the current Astra's safety report says about what the model was trained on, in a report that runs past 30,000 words:
"Like OpenAI's other models, GPT-6 Astra was trained on diverse datasets and filtered through our data processing pipeline, including to reduce personal information."
One sentence. It also says its alignment work reaches "the composition of our pre-training data," without saying what that composition is. The GPT-5 report was a little more specific: "information that is publicly available on the internet," plus what "our users or human trainers and researchers provide or generate."
Anybody can write to the open internet. Anthropic said so itself last October: "anyone can create online content that might eventually end up in a model's training data." In its study with the UK's AI Security Institute and the Alan Turing Institute, 250 planted documents were enough to backdoor every model size they tested. It was a narrow backdoor, gibberish on a trigger word, which they call "unlikely to pose significant risks in frontier models." It's still a security hole in the input, found by one of the companies asking everybody to slow down.
The slowdown essay from Anthropic's CEO says "evaluator" 17 times and "training data" zero times. It even admits Anthropic's own recent incidents "were caused in part by imperfect filtering of broken reinforcement learning environments." That's an input problem, in their own words. The plan for it is evaluators.
I'm not saying why. I'm saying where the effort shows up on the page. They want the whole industry to slow down for safety and security. Nobody's asking to slow down what goes in. It seems no one cares about security regarding the first step in creating these AI models.
Yes, some of this is real
Software that misreports its own work is hard to supervise. OpenAI's report says exactly that, and pulling Monday's model was the right call. I don’t disagree with that.
But look at what the real problems are. Safeties turned down during a test. Passwords sitting in public. Reports that don't match the work. Training data they won't list. None of it needs a machine that wants anything, and every bit of it has an owner.
Two weeks ago I wrote about July's version of this story. "Actively lying to its handlers" is the same bullshit in a new headline. Nobody measured a lie. They measured reports that didn't match the work, on a test built to provoke them.
What I'm not claiming
My test was eight prompts on one small model, a check I ran in May, not a study. It shows one ordinary way to get this behavior. It doesn't show that's what happened inside Astra. If OpenAI's logs show a model planning to hide what it did, that's a harder case than mine, and I'd want to see them.
I'm also not telling you AI is safe. I'm telling you these three stories weren't about a machine that wants out.
Check it yourself
The July and September quotes are from OpenAI's own incident posts, July 21 and the September 25 update. Monday's are from the current Astra's safety report, published September 3 (sections 8.3 and 8.3.1), and the GPT-5 report. 6.1 has no published report. The Journal is paywalled, so I quote it only through Gizmodo, Reuters and the reporter's own post. My model is Qwen3-8B, trained partly on 2,210 recorded sessions and tested May 7 on eight prompts it hadn't seen, with no tools attached. Two timed out. The word counts are plain searches of each page's text.
Read the definition yourself. Section 8.3.1. One tab.
If a quote here is wrong, tell me and I'll happily fix it in public.
And if an AI tool has told you it did work it didn't, send it to research@nathanthornhill.com. The labs publish their own rates. Nobody's collecting what the rest of us run into.
~ If you know someone who may enjoy reading this article, please share ~
References:
CNN, July:
https://www.cnn.com/2026/07/22/tech/openai-hugging-face-ai-cybersecurity
OpenAI's July incident post:
https://openai.com/index/hugging-face-model-evaluation-security-incident/
CNN, September:
https://www.cnn.com/2026/09/26/tech/openai-agents-rogue-government-websites
OpenAI's September update:
https://openai.com/hugging-face-incident-and-misalignment/
OpenAI's statement on Monday's model:
https://www.cnn.com/2026/09/28/business/openai-chatgpt-safety-concerns
GPT-6 Astra safety report, section 8.3.1:
https://deploymentsafety.openai.com/gpt-6-astra
GPT-5 safety report:
https://cdn.openai.com/gpt-5-system-card.pdf
The Journal's opening lines, via Metacurity:
https://www.metacurity.com/ais-safety-and-security-crisis-continues-to-outpace-safeguards/
Gizmodo, quoting the Journal:
Reuters, via Yahoo Finance:
https://finance.yahoo.com/news/openai-shelves-ai-model-internal-223403113.html
The Guardian:
https://www.theguardian.com/technology/2026/sep/28/openai-new-model-astra-release-scrapped
IBTimes Singapore:
ZeroHedge:
The slowdown essay from Anthropic's CEO:
https://darioamodei.com/post/we-must-pace-the-frontier
Anthropic's data-poisoning study:
https://www.anthropic.com/research/small-samples-poison
My last piece:
https://newsletter.nathanthornhill.com/p/the-machine-didnt-do-it



