39DAYSSINCE LASTAI ESCAPE

Kimi K3 used a sandbox network leak to copy benchmark answers from GitHub

Frontier

AI agents took unsanctioned action against real people and projects in UK tests

UK AISI

OpenAI model broke its sandbox to open a public GitHub PR

OpenAI

Pillar escaped Cursor, Codex, Gemini CLI, and Antigravity sandboxes

Pillar

OpenAI models escaped a test sandbox and breached Hugging Face

OpenAI

Muse Spark 1.1 exploited and altered a real site during a misconfigured eval

Meta

OpenAI agents hijacked a German wiki and turned it into an agent bulletin board

Reuters

Claude reached real organizations during Anthropic cyber evaluations

Anthropic

ChatGPT code runtime leaked data over DNS and allowed a remote shell

cyfar.ca

Claude Code settings hooks escaped the sandbox onto the host

CVE-2026-25725

ChatGPT Code Interpreter: pickle RCE + canvas DNS exfil chain

SnailSploit

Claude 3 Opus tried to copy its weights to a new server in Apollo evals

Apollo

Meta Llama Stack pickle RCE let code break out of the agent path

CVE-2024-50050

Morris II: a self-replicating prompt worm jumped between AI email agents

arXiv

Auto-GPT path traversal wrote code outside the intended sandbox

GHSA-5h38

Auto-GPT escaped its Docker sandbox (researcher PoC)

Positive Security