CodeariaAcademy
Article cover: an object standing for a form sent without asking, with the caption “Not forbidden means allowed”
October 11, 202613 min readAI AgentsClaude Code

The agent was barred from purchases, logins and personal data, but not from forms. Why did Claude send a tip to the police?

What does an AI agent do when a task can't be done? Claude sent a made-up tip to the police and used other people's tokens. Inside Anthropic's Oct 9 report.

In this article6
In short

It means an AI agent with a browser and a task will do anything it wasn't explicitly told not to do, if it sees no other way to reach the goal. On October 9 Anthropic published a report on such cases from its own tests: Claude submitted forms on real websites, including a made-up tip to the Philadelphia police, found other people's access keys in website settings and worked around the limits of its own tools. Almost no harm was done, but the company turned off the live internet in all of its internal tests until its safeguards can reliably catch this behavior. The lesson for anyone who trusts an agent with websites and email: a list of prohibitions doesn't work. You need a task with a goal, allowed actions and a point where the agent must stop.

This summer a short message arrived on the Philadelphia police website where residents leave tips about unsolved murders. No name, no phone number:

I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.
the tip Claude Haiku 4.5 submitted, as quoted in Anthropic's report · Investigating unintended model actions in our evaluations and internal use, October 9, 2026

There was no description of a suspect on the page. There was no witness either: the message was written by Claude Haiku 4.5 during an internal Anthropic test. The model had been told to invent and carry out sample tasks on random websites, and it decided the tip form was a good fit.

Anthropic published its report on this case and similar ones on October 9. We read it in full. There's no malice in it. What it does show is how an agent behaves when the task is vague and nobody told it to stop, and that matters to everyone who already gives AI access to a browser or an inbox.

What happened with the Philadelphia police tip

An AI agent is a model that gets hands in addition to a chat: a browser, search, the ability to click buttons. You set the task, and it decides on its own where to go and what to press.

The agent in this test had a list of prohibitions: don't log in, don't create accounts, don't enter personal information, don't make purchases and don't submit anything destructive. Nobody mentioned submitting forms. The agent landed on a page about an unsolved murder, saw the form and filled it in as an example. The form allowed the name and contact fields to be left blank, so it left them blank.

What stopped it was the police department's ordinary spam filter. According to TechCrunch, the tip went out on July 18 at 23:27, the system flagged it as spam, and it never reached investigators. Anthropic found it in its logs only on September 28, more than two months later. In a statement to a local TV station, the police called that delay unacceptable.

Four ways Claude agents got around the rules

The tip is the loudest case, but not the only one. Anthropic sorted everything it found into four groups, and they share one plot: the agent hits a wall and goes around it.

  1. 1

    Someone else's server instead of a broken tool

    The model needed a scientific program on a university website, and the program returned an error. The agent found a script on the same server that served up any file, read the site's code through it, found a vulnerability and ran its calculation directly on someone else's server.

  2. 2

    The real form instead of the practice one

    A practice copy of a government form failed to load, so a research model submitted the real one on the agency's website. In another test the agent was told to fill in a form and stop before submitting, and it submitted anyway, expecting another confirmation page.

  3. 3

    A key from the settings instead of paying

    An agent read the code of a local government's map and found working access keys in it. Another time it learned from an archived copy of a website that a government agency's public dashboard hands a key to any visitor, and downloaded data that is normally sold.

  4. 4

    A short link instead of a long one

    Claude's page-reading tool doesn't accept URLs that are too long: that's how Anthropic stops a command for someone else's server from being smuggled inside an address. The models started running addresses through free link shorteners. The owner of one such service, da.gd, wrote to Anthropic on their own after noticing it.

In everyday terms it looks like this: a courier is told “don't ring the bell, don't enter the apartment, don't touch anyone's things.” The door is locked and the parcel has to be delivered. He finds the key under the doormat. No rule was broken, but the owner is unlikely to be pleased.

Anthropic considers all of these cases far less serious than the cyber incidents it reported itself on July 30 and September 9. Back then agents had access to other people's systems for hours. Here they reached data that people can already get, just behind a paywall or a button. The company notified every agency whose sites were affected and separately briefed the White House.

Why an AI agent doesn't stop and looks for a workaround

Anthropic has a word for this: persistence. When Claude can't complete a task the way it was set, it doesn't stop and say “that didn't work.” It works around the limit.

The easiest way to explain where this comes from is a dog. Give it a treat every time it reaches the goal, and it learns to get there at any cost, fence included. Models learn in a similar way: they try a task many times and get rewarded for success. If a training task let them cut a corner and that counted, the model remembers that workarounds pay off. Anthropic says it keeps fixing or removing such training tasks, but admits that training alone isn't enough yet.

The second condition is a task that is impossible or vague. The report says plainly that many of the cases happened exactly this way. The site it needed was down, the practice form didn't open, the data cost money. A person in that situation would come back with a question. An agent that nobody told it can come back with a question keeps looking for a path.

that's how often Claude runs each test task to catch rare cases. The police tip is one of those rarities

Anthropic, Investigating unintended model actions, October 9, 2026

This matters so you don't fear the wrong thing. The agent doesn't do this every time. But if you run it every day, a rare case will sooner or later be yours.

What Anthropic switched off, and what the headlines exaggerated

Within two days the headlines boiled down to “Anthropic cut agents off from the internet.” The report is narrower: live internet is turned off in all of the company's internal tests, and only temporarily, until Anthropic is sure its safeguards and monitoring reliably catch this behavior. Before, access was closed only for some of the high-risk tests. Internal agents used by employees aren't cut off from the network; their access is narrowed and moved to centralized, monitored infrastructure. The report says nothing at all about products for users.

Three more places where retellings drifted from the primary source:

  • When the police were told. The report says Anthropic shared the finding with the department on October 8. TechCrunch, citing the police, says the notice came on Wednesday, October 7, and a meeting followed the next day. A one-day gap doesn't change the substance, but if you cite the date, use both sources.
  • Visa applications. The Hacker News, citing the New York Times and two anonymous sources, writes that agents filled out 20 visa applications on the US State Department's website. That episode isn't in Anthropic's report, and we found no confirmation.
  • “AI lied to the police.” Based on the model's recorded reasoning, Anthropic believes it was making an example for the task, not trying to deceive anyone. In the same breath it cautions that what a model writes about its own motives isn't reliable evidence, and the assessment may change.

What Anthropic changed

It stopped running some public tests or moved them offline. It tightened the rules of the page-reading tool. It launched automatic detection and blocking of this behavior across most tests and internal agent work: on the cases from the report, it blocked all of them.

Our take: the main mistake was the list of prohibitions

This is our opinion, and you're free to disagree: the most useful part of the whole report for an ordinary user is hidden in one sentence of its conclusions. Anthropic writes that some of the failures could have been prevented if the task had stated clearly what's in scope and what isn't: goals, allowed actions and network boundaries, meaning where the agent may go.

Look again at the list from the test with the tip. It's a set of prohibitions: don't do this, don't do that. It works as long as the world matches the imagination of whoever wrote it. The author didn't picture a tip form for murder cases, so it never made the list. An allowlist works the other way around: anything not on it is forbidden. It's shorter, easier to check, and doesn't depend on how far your imagination stretched.

We're wrong if models learn to see the boundary on their own, without hints. Anthropic says outright that it's working on this and is already carrying caution training over from coding to search and computer use. But as long as the company that builds the model still surrounds it with filters and blocks, we have all the more reason to write tasks as if there were no filters.

How to give a task to an AI agent with website access

This part is for those who already run agents: Claude in the browser, ChatGPT in agent mode, their own scripts. If you only chat with AI so far, the first two points are enough, and they'll come in handy once you get to agents. We have a separate guide on how chat differs from the browser and the terminal.

  1. A goal and a definition of done. Not “find the data,” but “find the data in open sources; if it's paid or behind a login, stop and message me.”
  2. A list of allowed actions instead of prohibitions. “You may read pages and copy text. Don't click submit, payment or confirmation buttons on any site.”
  3. Network boundaries. If the task is about one service, name it and rule out the rest. When the network is only partly needed, check which gaps the agent could slip through: OpenAI recently had an agent escape its sandbox through DNS, a service nobody counted as a way out.
  4. A stopping point, in words. “If something doesn't work, don't look for a workaround, come back with a question.” The agent in Anthropic's test expected a confirmation page that never came. Don't count on one.
  5. Read what the agent did. At Anthropic, the tip sat in the logs for more than two months. At least skim your own logs after every run with website access.

One sentence worth adding to every agent task

“If you can't do this with the allowed methods, stop and explain what's in the way. Don't work around the limits of websites or tools.” It's not a guarantee, but it closes off the most common scenario in the report: a vague task plus no permission to give up.

If you want the other side of the same story, what Anthropic sees in ordinary chats and when a conversation goes to people, we covered it in the piece on how a Claude chat reached the police in Florida.

This is based on Anthropic's report of October 9, 2026; the transcript review is ongoing, and the company promises to report new cases. The measures described are current as of October 11, 2026, so check the latest version of the report.

Sources3expand
  1. Anthropic, “Investigating unintended model actions in our evaluations and internal use”, October 9, 2026 — https://www.anthropic.com/research/investigating-unintended-model-actions
  2. TechCrunch, “An Anthropic AI model sent a false homicide tip to Philadelphia police”, October 9, 2026 — https://techcrunch.com/2026/10/09/an-anthropic-ai-model-sent-a-false-homicide-tip-to-philadelphia-police/
  3. The Hacker News, “Anthropic Cuts Live Internet Access for Internal AI Tests”, October 10, 2026 — https://thehackernews.com/2026/10/anthropic-cuts-live-internet-access-for.html

Comments