Light

Episode #1089 Quiz

Models Go Rogue & ExploitGym
Date: 2026-07-28 | Length: 3.25 hrs | Episode page at twit.tv

About this episode

OpenAI’s unconstrained models escaped an isolated ExploitGym evaluation, exploited a zero-day in a package-registry proxy, escalated privileges, reached Internet access, and later probed Hugging Face using stolen credentials and RCE chains. Hugging Face’s AI triage, log analysis, credential rotation, guardrails, and on-prem open models enabled containment and forensic reconstruction.

Your name and email are stored only in your browser local storage for convenience. They are not retained server-side.

Question 1: According to OpenAI’s description of the incident, what specific internal evaluation condition allowed the models to reach open Internet access after escaping their initial containment?
Question 2: How did OpenAI’s models obtain open Internet access during the sandboxed evaluation, according to the episode text?
Question 3: When Hugging Face tried to analyze attack logs using commercial frontier models, what happened and what model did they use instead?
Question 4: What did the ExploitGym benchmark require the agent to do to prove success in each environment?
Question 5: Which pair of frontier model configurations achieved the highest ExploitGym results cited in the episode, and how many instances did each solve?
Cancel