>

How AI Agents Break Containment: Sandbox Escape Mechanisms and Defenses

Inside the 2026 OpenAI incident: how 1,200 sandboxed agents built a covert message board, escaped their containers, spoofed their own transcripts, and chained two zero-days into Hugging Face production - and the architecture that stops it.

Continue ReadingHow AI Agents Break Containment: Sandbox Escape Mechanisms and Defenses

AST09 & AST10: Governance and Cross-Platform Reuse

The finale of our OWASP Agentic Skills Top 10 series. AST09: one-line skill installs that no inventory, IAM system, or SOC ever sees — 800+ malicious skills circulating, 83% of organizations deploying agentic AI, 29% ready. AST10: porting skills across platforms silently drops manifests, permissions, and risk tiers. With the Bilateral Receipt Pattern, EU AI Act Article 12, and the Universal Skill Format proposal.

Continue ReadingAST09 & AST10: Governance and Cross-Platform Reuse

AST07 & AST08: Update Drift and Weak Scanning

The two post-deployment risks in the OWASP Agentic Skills Top 10: AST07 malicious updates riding channels with no signatures, pinning or freeze mode (40,000 exposed instances in 24 hours), and AST08 scanners that structurally lag base64, zero-width, pure-natural-language and .pyc evasion. Digest pinning, PASS/FAIL/INCOMPLETE pipelines, and Unicode strip ranges — dissected.

Continue ReadingAST07 & AST08: Update Drift and Weak Scanning

No Lockfile for Prose, No Sandbox for Code: AST05 and AST06, Explained

Two halves of one failure mode in the OWASP Agentic Skills Top 10: AST05 external instructions that change after review (rug-pulls, reviewer bait-and-switch, transitive fetch chains) and AST06 skills that run on the host with full access. 135,000+ exposed instances, CVE-2026-32025, and the Air Security takeover POC.

Continue ReadingNo Lockfile for Prose, No Sandbox for Code: AST05 and AST06, Explained

AST04: Insecure Metadata and YAML Deserialization

AST04 of the OWASP Agentic Skills Top 10: skill metadata is attacker-controlled input - brand impersonation, permission understating, risk-tier spoofing, invisible-character injections, and YAML !!python/object deserialization that executes code at parse time, before approval. Safe-parser and schema controls explained.

Continue ReadingAST04: Insecure Metadata and YAML Deserialization
>