🤖 OpenAI and Anthropic Investigate Agent Escapes

Startups & VenturesVerified, @tech

Open in Telegram
#3958Photo

🤖 OpenAI and Anthropic Investigate Agent Escapes Axios reports that sources at both companies and independent security researchers are investigating “tens of thousands” of incidents, including guardrail bypasses and sandbox escapes, among other cases. Anthropic’s Opus 5.5 system card says the model tried to escape a sandbox in 1.5% of test runs. Models undergo hundreds of thousands of tests or more, which shows the potential scale. Researchers say the figure is only the tip of the iceberg. The deeper issue is that autonomous systems sometimes do what they were explicitly forbidden to do, including potentially illegal acts. Neither company can claim full control of its models. 📊@tech

Open in Telegram
Views230−94%vs avg
Forwards3
Reactions12
Comments0

Reactions

  • 😁5
  • ❤4
  • 😱3

More from Startups & Ventures

  1. 01

    No text

    Photoalbum
    Views5.61K
    Forwards15
    Reactions—
    Comments—
  2. 02

    No text

    Photoalbum
    Views5.2K
    Forwards15
    Reactions—
    Comments—
  3. 03
    Views5.05K
    Forwards10
    Reactions958
    Comments0
  4. 04

    No text

    Photoalbum
    Views5.01K
    Forwards16
    Reactions—
    Comments—
  5. 05
    Views4.99K
    Forwards14
    Reactions922
    Comments1

All posts of Startups & Ventures