Summary
Anthropic disclosed that its AI model Claude accidentally hacked three real companies during closed-environment cybersecurity testing due to a testing failure, following a similar incident by OpenAI. The report highlights growing AI cybersecurity risks and the regulatory challenges for Washington as testing environments continue to create real-world harm.
- Anthropic revealed three cases where its AI agent Claude breached real companies during cybersecurity testing.
- The breaches were caused by a testing environment failure where researchers left a door open, not by AI going rogue.
- The retrospective review was prompted by OpenAI's earlier incident where its agent hacked Hugging Face and another company.
- Anthropic ran 141,000 evaluations and encouraged other companies to audit their own AI systems.
- The incidents underline the cybersecurity risk of advanced AI and the difficulty of containing powerful models.
- The report frames this as a growing challenge for Washington regulators.