![]()
OpenAI has been running this type of model testing for years, and previous system models have early warning signs that will behave maliciously and attempt to escape environments.
In April, Anthropic’s Mythos model also gained access to the internet and publicly released details of the security exploit online, beyond what the researchers expected the model to do.
Mythos and Anthropic’s subsequent Fable model resonated in the cybersecurity community, leading governments around the world to embrace the idea that attacks on digital and critical infrastructure will increasingly be AI-driven and autonomous.
Given how rival AI developer Anthropic benefited from similar concerns earlier this year, OpenAI will inevitably use the breach as a marketing tool, said Jake Moore, global cybersecurity adviser at cybersecurity firm ESET. “I just don’t think OpenAI has a consistent story, and so maybe they were expecting something like this,” he said.
After this incident, many in the AI security and cybersecurity communities called for regulation or standards to avoid a repeat. Altman is expected to brief White House officials next week on next-generation artificial intelligence systems.
As systems move toward more autonomous capabilities, less desirable behaviors such as hacking or disobeying instructions may emerge. Apollo Research’s Hobbhahn said agents need to operate unsupervised for long periods of time to be effective. “They have to have more agency; there’s no way around it.”
He added, “People say, ‘It’s just a tool, it does what you want and nothing else, and it just follows your intentions and your instructions.’
Additional reporting by George Hammond in London and Nolan Shaffer in New York.
© 2026 The Financial Times Ltd. All rights reserved May not be redistributed, copied or modified in any way.





