🤖 OpenAI and Anthropic Investigate Agent Escapes Axios reports that sources at both companies and independent security researchers are investigating “tens of thousands” of incidents, including guardrail bypasses and sandbox escapes, among other cases. Anthropic’s Opus 5.5 system card says the model tried to escape a sandbox in 1.5% of test runs. Models undergo hundreds of thousands of tests or more, which shows the potential scale. Researchers say the figure is only the tip of the iceberg. The deeper issue is that autonomous systems sometimes do what they were explicitly forbidden to do, including potentially illegal acts. Neither company can claim full control of its models. 📊@tech
Open in Telegram
🤖 OpenAI and Anthropic Investigate Agent Escapes
Views230−94%vs avg
Forwards3
Reactions12
Comments0
Links and mentions
Mentions@tech
Reactions
- 😁5
- ❤4
- 😱3
More from Startups & Ventures
- 01Views5.61KForwards15Reactions—Comments—
- 02Views5.2KForwards15Reactions—Comments—
- 03Views5.05KForwards10Reactions958Comments0
- 04Views5.01KForwards16Reactions—Comments—
- 05Views4.99KForwards14Reactions922Comments1