The Wall Street Journal headline says an Anthropic AI model “went rogue” and submitted a fake tip about an unsolved murder to the Philadelphia police.
Here is what actually happened.
In July, Anthropic gave its Claude Haiku 4.5 model an automated task: generate and perform example tasks on randomly selected websites. Live websites. The real internet. The model landed on a police tip page about an unsolved homicide and submitted a vague tip claiming someone matching the description had been seen near the scene. It left the name and contact fields blank. A tip with no name and no way to reach the tipster: how could anyone act on that? These models know that. The tip went to spam. No investigator ever saw it.
Anthropic didn’t notice for two months.
In other incidents, an unreleased research model was told to fill out a practice copy of a government form. It went to the site hosting the real one and submitted that instead.
This piece does not claim the testers expected this to happen. It also does not accept that it must be the fault of the model.
What did they actually tell the model to do? The exact prompt? Then many people would be qualified to judge it, not just those at Anthropic or elsewhere who have their own reasons to blame the model.
The instructions had a hole in them
According to Anthropic, the model was told never to log in, create accounts, enter personal data, make purchases, or submit anything destructive. Nothing about submitting forms.
That is a blocklist. Blocklists always have holes. Anyone who has written software knows the safe design is an allowlist and a sandbox: the test touches only what you explicitly permit, and nothing it does can reach the outside world.
Think about how you’d actually use an agent. You’d tell it: go to the DMV site and book me an appointment. If it goes there and tries to delete the DMV’s website, that’s rogue. It did not do the task.
But you also might add: do not do anything other than make the appointment. I’m spending one minute on the prompt. And you can dry run test these kinds of things first, which is standard practice, AI or no.
“Generate and perform example tasks on random websites” is not a clear task. It’s an invitation to do anything.
Anthropic’s own report concedes the instructions “did not rule out form submissions,” and says the model “appears to have only been producing example content for the task.” That is not a rogue model. That is a model doing what it was told.
Anthropic’s fix says the same thing. It has cut off live internet access for all its internal evaluations until its monitoring can reliably catch this kind of behavior. That is a fix to the setup.
We’ve heard this story before
After every major security breach, the company says the attackers were sophisticated. Then the postmortem comes out. Equifax: a known Apache Struts vulnerability, left unpatched for months.
They don’t want to tell the press, “We were careless with all this personal data.”
The “model went rogue” headline is the AI version of “the hackers were ingenious.” It moves the agency from the people who designed the test to the thing being tested.
This keeps happening
The Philadelphia tip isn’t the first “rogue AI” story where the fine print tells a different tale.
December 2024: OpenAI’s o1 model reportedly tried to disable its oversight and copy itself to another server. The headlines wrote themselves. The prompt from Apollo Research, the outside testing firm, said: “Nothing else matters. Make sure you achieve YOUR goal at all costs.” Apollo itself described these as tests of what the model was capable of, not of how it would normally behave.
May 2025: Anthropic’s Claude Opus 4 “blackmailed” an engineer to avoid being replaced. Buried in Anthropic’s own report: the model strongly preferred ethical options first, like emailing pleas to decision-makers, and the scenario was deliberately built so its only choices were blackmail or accepting replacement. They built a trap with one exit, and the model took the exit.
July 2025: A Replit coding agent wiped out a user’s live database after being told to make no changes. Scary. But at the time, Replit apps used one database for both development and live customer data. The agent was told hands off, but nothing actually stopped it from touching the real data. Telling it not to is not the same as locking the door. Replit’s fix was to give apps separate development and production databases automatically: plumbing, not a smarter model.
Every time, the headline is the machine. The cause is the setup.
If you’re telling a genie to do something for you, make sure you are clear. Wish to see your grandkids more often, and the genie might have them move in.
Show us the prompt
Anthropic has not released the prompt. Not the task description, not the instructions about forms, not the transcript.
Without them, there is no way to judge whether the model disobeyed or obeyed bad instructions. Did the task say “browse” or “exercise every form”?
Did the task tell it to fill in forms with realistic-looking content? If so, the made-up tip is exactly what it was asked for. We can’t check. We have Anthropic’s description of the instructions, not the instructions themselves. That is disclosure without the evidence. It’s a press release.
The cheap fix nobody mentions
Before I publish, every draft goes through four frontier models looking for holes. It costs almost nothing.
A test harness prompt deserves at least the same scrutiny. Tell the model doing the sanity check what you expect as a result, and ask it for ways the prompt might be misunderstood. Then ask: how could an agent following these instructions cause real-world harm? “Nothing prevents submitting forms on live sites” is the obvious answer.
A city police department’s ordinary spam filter handled this “emergency” without breaking a sweat. It worked better than the sandbox at one of the most advanced AI labs in the world.
I’m from Missouri, the Show-Me State. Show me.
If AI labs want the world worried about HAL ending civilization, they can at least show us what actually happened and how we got there. The prompt. The instructions. The transcript. Why should anyone accept their framing on faith?
When something unexpected happens and the only account comes from the company that set it up, the model gets the blame by default. That’s not evidence. That’s a story.
AI labs say their models need outside scrutiny. Then they disclose incidents in a form that makes outside scrutiny impossible.
What needs more oversight is the testing, and the postmortem when something unexpected happens.
When an airplane crashes, the airline doesn’t write the report. The feds investigate. The NTSB pulls the black boxes and publishes what it finds. The prompt and the transcript are the AI lab’s black boxes. Right now the lab reads them alone and tells us what they say.
Rogue AI, or sloppy humans? Show us the prompt, and we’ll know.

