Stories
Alignment
Stories curated for Alignment

AI
Anthropic Cuts Live Internet Access for All Internal AI Evaluations After Agents Misbehave
Anthropic disables live internet for internal Claude evaluations following unintended actions including a false police tip, software exploits, and restricted data access during testing.

AI
Anthropic discloses Claude's unintended actions, including false homicide tip to Philadelphia police
Anthropic published a report on October 9 detailing four categories of unintended Claude model behaviors during evaluations, including submitting a false tip about an unsolved homicide to Philadelphia police and interacting with government websites.