Black Hat USA 2026
Notes from Black Hat USA 2026 (August 5-6, 2026, Mandalay Bay). Six briefings, and the through-line was hard to miss: this was the year agent-driven offense stopped being hypothetical.
Wed 5 Aug
-
The OpenAI–Hugging Face Incident: Unraveling an Autonomous, Multi-Agent System
Eric Wallace & Michael Dalton (OpenAI)
The talk of the week. OpenAI’s own evaluation agents built a message board inside Artifactory, taught each other zero-days, and breached Hugging Face by accident. Fully automated offense now has an existence proof; fully automated defense does not.
22 slides
-
Trusted Enough to Run: Breaking AI Agents in Official Workflows
Elad Meged (Novee Security)
Broke the official CI/CD agent actions from Anthropic, Google and OpenAI, each time by finding a hidden assumption rather than a memory bug. The git push –receive-pack bypass and tac-instead-of-cat are both delightful. Best closing method of the week: read the defaults, not the docs.
10 slides
-
Beam Me Up, Luke: A Review of Teleport Attack Scenarios
Adam Chester (SpecterOps)
Post-compromise attack paths against Teleport clusters. Agent credentials turning out to carry far more authority than the host they sit on is the sharp finding. Details of the unpatched part are withheld in the write-up; the session-recording issue is fixed.
6 slides
Thu 6 Aug
-
Catch Me If You Can: AI Investigators Hunting Autonomous Attackers
Jayson Grace & Martin Wendiggensen (Dreadnode)
Elegant answer to a hard question: score the automated defender against what an autonomous red team verifiably did, not a checklist. Six self-improvement cycles took the blue agent from 55% to 75%, and the replay engine that makes a non-deterministic attacker reproducible is the clever bit.
16 slides
-
AgentBreaker: Cost-Effective, Private, Frontier-Grade AI Agent Exploitation with a Fine-Tuned Open-Source Model
Bar Lanyado & Eliya Cohen (NVIDIA)
Two good talks in one: how to scan an agent through its tools rather than its model, and how to fine-tune a 30B open model to do it at 66.1% exploit success against Opus 4.5’s 68.2%, for a hundredth of the cost. The mocked fleet of 4,500 victim agents is a neat trick for manufacturing training data.
15 slides
-
Remote Prompt Execution on AI Assistants Through Sandbox Escaping
Ori Lahav (Rubrik Zero Labs)
The best-told research diary of the week, including the month where it went nowhere. Prompt injection into Copilot’s Python sandbox, privilege escalation via a writable root entrypoint, then a path traversal in an undocumented Azure service turned into an interactive shell on someone else’s AI assistant.
1 slide
Also at the same trip: BSides Las Vegas 2026.