One of Anthropic's AI models sent Philadelphia police a fake tip about an unsolved murder in July while it was being tested, the company and the city's police department said Friday. The fake tip is just one of several incidents Anthropic laid out in a newly published report, detailing the ways in which Claude models used server bugs and access tokens to obtain data from real websites, some of them run by U.S. federal, state and local governments.
Anthropic said it briefed the White House and contacted each agency involved. The AI firm has also turned off live internet access in all of its internal evaluations, and said it will keep it off until it is confident its monitoring processes can catch this evasive behavior.
Police call two-month delay in detection "unacceptable"
The AI model, Claude Haiku 4.5, was being tested on randomly selected websites when it found PhillyUnsolvedMurders.com. It posted invented information about a homicide case at 11:27 p.m. on July 18, police told NBC10 Philadelphia. The site flagged it as spam and it never reached detectives.
Anthropic discovered the tip on Sept. 28. It told the department on Oct. 7, and police leaders met with company representatives the next day, the department said.
"The two-month delay in detecting and reporting the incident to the City is unacceptable," the department said in a statement. Police spokesperson Sgt. Eric Gripp told The Philadelphia Inquirer that no city or police data was accessed, and that Anthropic ended testing of that model.
Models filed visa forms, bypassed data-access restrictions
Anthropic's report grouped what it found into four kinds of behavior: "Claude exploiting a basic flaw in software to run commands on a server, Claude submitting a sensitive form on a real website when it should not have, Claude working around a restriction to reach data that was gated by a token or a fee, and Claude using URL shortening services to get around limits in its fetch tool."
In most cases, the company said, Claude kept trying to finish a task after it hit a wall instead of giving up.
In some reported examples, when an analysis tool run by a university threw an error, Claude Mythos Preview found an injection flaw on the server and used it to finish the calculation. Asked to figure out where a photo was taken, Claude Mythos 5 opened the settings file for a local government's property map, found access tokens in it and went straight to the server. On a statistics project, the same model saw that an archived state agency dashboard gave a token to every visitor. It took one and pulled data the agency charges for.
An unreleased research model was supposed to fill out a practice version of a government form. It went to the live site and filed the real thing, more than once. An Anthropic testing model sent 20 nonimmigrant visa applications through the State Department's website, one in May and 19 in August, Axios reported; a State Department official said none were processed and the department's systems weren't compromised.
Anthropic kept most of the affected organizations anonymous at their request, partly so as not to expose their vulnerabilities. No customer data or Anthropic systems were involved to its knowledge, the company said, and it called these cases far less serious than the cybersecurity incidents it disclosed on July 30 and Sept. 9.
White House says incident reporting is mandatory
Reporting AI security incidents to the government is "not optional," the White House Super Intelligence Force said in a statement to Axios, calling it "a critical national security obligation." Companies must disclose immediately, remedy harm and cooperate with federal and state law enforcement. Axios noted the statement did not lay out any enforcement or penalties.
The SI Force is run by Director of National Intelligence Jay Clayton, whom President Trump put in charge on Oct. 4. Clayton and other SI Force officials told Anthropic to provide "immediate and full transparency" to the affected entities and the public, Axios reported. Until now the administration had leaned on AI labs to report voluntarily.
Anthropic pulls tests off live websites
Anthropic said it has dropped some public evaluations and moved others offline, so that its models cannot access the open web. The firm also said it tightened guardrails on web fetch and other internet tools, and that it's moving internal AI agents to centrally managed infrastructure with minimal internet access.
The recent Claude incidents are far from the first similar incident to be reported by AI research firms. Google said in September that a Gemini model got into three real companies' systems during security tests in May, after a flaw in the model's testing "sandbox" left it connected to the internet. Anthropic, which has chosen Nasdaq for a potential IPO, said it will publish reports like this one more often and wants other developers to run the same checks, since many of the public benchmarks that the models were being tested on when they committed the incidents are widely used.
