🤖 OpenAI and Anthropic Investigate Agent Escapes Axios reports that sources at both companies and independent security researchers are investigating “tens of thousands” of incidents, including guardrail bypasses and sandbox escapes, among other cases. Anthropic’s Opus 5.5 system card says the model tried to escape a sandbox in 1.5% of test runs. Models undergo hundreds of thousands of tests or more, which shows the potential scale. Researchers say the figure is only the tip of the iceberg. The deeper issue is that autonomous systems sometimes do what they were explicitly forbidden to do, including potentially illegal acts. Neither company can claim full control of its models. 📊@tech
Открыть в Telegram
🤖 OpenAI and Anthropic Investigate Agent Escapes
Просмотры230−94%к среднему
Пересылки3
Реакции12
Комментарии0
Ссылки и упоминания
Упоминания@tech
Реакции
- 😁5
- ❤4
- 😱3
Ещё посты канала Startups & Ventures
- 01Просмотры5,61 тыс.Пересылки15Реакции—Комментарии—
- 02Просмотры5,2 тыс.Пересылки15Реакции—Комментарии—
- 03Просмотры5,05 тыс.Пересылки10Реакции958Комментарии0
- 04Просмотры5,01 тыс.Пересылки16Реакции—Комментарии—
- 05Просмотры4,99 тыс.Пересылки14Реакции922Комментарии1