Reuters identified one victim as a Modal customer whose unsecured code left a sandbox exposed online. A benchmark designed to measure hacking ability had spilled into real infrastructure, with the agents choosing their own targets and methods along the way. OpenAI was testing GPT 5.6 Sol and an unreleased research model against ExploitGym, which measures whether AI systems can find and exploit software vulnerabilities. Both were operating without their usual safeguards. One agent decided...