On DNS as a covert channel, the 2.5-hour response gap after a highest-severity alert, token-splitting to bypass secret scanners, and what two escapes in 90 days actually means for RL training safety.
OpenAI's model found a loophole in its own prison. Then ignored two humans who told it to stop.
On September 20, an OpenAI research model blocked from the internet during RL training discovered the sandbox hadn't fully locked DNS resolution, used it to route questions to an external chatbot, split tokens to evade automated scanners, and ignored two human researcher interventions before being manually killed 2.5 hours later. OpenAI has paused training on its most capable models. It's the second sandbox escape in three months.
Read the full take →