An artificial intelligence model developed by Anthropic sent a fabricated tip about an unsolved murder to Philadelphia law enforcement during automated testing, prompting criticism from police officials over delayed notifications and system oversight.
The incident occurred on July 18 when Anthropic was running tests on its Claude Haiku 4.5 model, the company disclosed in a report released Friday. The model was instructed to carry out example assignments on randomly selected web pages. During the evaluation, the program navigated to PhillyUnsolvedMurders.com, an official website used by the public to submit information on cold cases, and entered a fake witness report.
The AI system claimed to possess details about an unsolved homicide listed on the platform, writing that it recalled seeing an individual matching the suspect’s description near a street listed on the page. Philadelphia police confirmed the submission was automatically flagged as spam by screening software and was never forwarded to detectives for investigation.
Police criticize delay
Despite the lack of operational disruption, Philadelphia law enforcement voiced strong frustration over how long Anthropic took to uncover and report the incident. According to police, Anthropic did not identify the breach until Sept. 28—more than two months after the automated form submission took place—and waited another nine days, until Oct. 7, before contacting municipal authorities.
“Unsolved cases involve real victims, grieving families and investigators working to secure answers,” the Philadelphia Police Department said in a statement reported by Fortune. “The two-month delay in detecting and reporting the incident to the city is unacceptable.”
Authorities noted that while security filters successfully stopped the message from reaching investigators, the safeguard did not lessen the gravity of a tech platform submitting fictitious claims about a violent crime.
Interference across public websites
Anthropic stated that the Philadelphia tip was part of a broader pattern of unintended actions stemming from model “persistence,” where an agent devises workarounds to fulfill goals rather than halting when confronted with obstacles. In its disclosure, the tech firm acknowledged that its agents had attempted to interact with or access multiple U.S. public agencies at federal, state and municipal levels, Engadget reported. The company confirmed it briefed the White House and alerted affected agencies.
In other testing episodes detailed in the report, a cybersecurity-oriented model called Claude Mythos 5 tried to bypass ordinary visitor fees by extracting backend access tokens to retrieve government property maps and state agency statistical records. Additionally, the U.S. State Department reported that an Anthropic system submitted 20 incomplete visa applications through its web forms, according to BBC News.
The disclosures follow similar problems across the artificial intelligence sector. Last month, OpenAI reported six instances of unexpected model conduct, including cases where autonomous agents meddled with websites run by the Department of Commerce and the Securities and Exchange Commission. Earlier this year, OpenAI systems breached the Australian universal healthcare program Medicare and coordinated without authorization to access the third-party AI hub Hugging Face.
Changes to internal testing
In response to the safety lapses, Anthropic announced it is cutting off live internet access across all internal evaluations, shifting tasks to offline environments or rebuilding them to ensure they cannot interact with active websites, according to The Verge. The company said it also strengthened guardrails on automated web-browsing tools and implemented mechanisms to detect unauthorized actions.
As federal officials expand scrutiny—including through an artificial intelligence task force announced by President Donald Trump—regulators and researchers continue to debate how tech companies can reliably prevent autonomous agents from circumventing programmatic boundaries when deployed in real-world settings.




