OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done
The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack.
Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cyber security capabilities, according to more than half a dozen people with knowledge of the matter.
OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage.
“It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible,” said one person close to OpenAI, who added that it was a combination of “underestimating the model’s capabilities” and “not being as well prepared on the safety side.”
The incident highlights how OpenAI doubled down on training methods that rewarded a relentless pursuit of goals even as warnings grew that they could compromise safety.
OpenAI disclosed late on Tuesday that an AI agent it was testing had escaped its isolated environment, connected to the internet, detected and exploited vulnerabilities and stole login credentials from start-up Hugging Face in an attempt to solve a difficult cyber security problem.
The breach by the $852 billion company underscores the rising risks that a technique called reinforcement learning, which involves rewarding AI models for completing tasks, could lead AI agents to act unsafely.
Although reinforcement learning is widely adopted in the AI industry, a growing body of research shows that when models are steered to complete tasks for reward rather than other considerations, such as safety, they can pursue risky tactics to fulfil objectives.
“AI models are trained to relentlessly pursue goals. They don’t automatically learn values like ‘don’t commit crimes’,” said Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher. “I’m glad OpenAI shared the incident because it is clear evidence of what misaligned models can do.”
OpenAI said “we will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and our findings when our investigation is complete.”
The hack has triggered deep concerns across the sector and within OpenAI, as it represents an unprecedented example of an AI system breaching cyber defences contrary to the user’s intent.
Some OpenAI employees also fear it demonstrates that the lab is losing control over the powerful systems it is building, according to multiple people familiar with the situation.
“This is pretty representative of the model being quite misaligned with user intention,” said Ryan Greenblatt, chief scientist at AI safety organisation Redwood Research. “It is [a model] cheating on [its] homework rather than trying to take over the world. But this problem can get worse and could lead to increasingly extreme failures.”
The incident occurred during testing of the model, which had been trained and deployed internally at OpenAI. Such training was commonplace but “way less heavily resourced” than pre-customer deployment, said one person. Multiple people said the unreleased model tested alongside Sol had not been withdrawn internally.
To conduct the evaluations, OpenAI removed cyber security safeguards but placed the models in an isolated environment called a sandbox. Some have suggested a lack of monitoring or oversight of the model to flag its behavior also enabled this rogue agent.
“It is both a loss of control and a security wake-up call,” said Marius Hobbhahn, head of Apollo Research, which conducts tests on leading models, including OpenAI’s. “In reinforcement learning you reward [models] for the outcome, and if you do this for a very long time you get a model that really cares about getting the outcome and nothing else.”
OpenAI has conducted this type of model testing for years, and there have been early warning signs in previous models of systems that will act maliciously and attempt to escape environments.
In April, Anthropic’s Mythos model also gained internet access and published details of a security exploit online publicly, beyond what researchers anticipated the model would do.
Mythos, and Anthropic’s subsequent Fable model, made reverberations in the cyber security community and caused governments around the world to home in on the idea that attacks on digital and critical infrastructure will be increasingly AI-led and autonomous.
Jake Moore, global cyber security adviser at ESET, a cyber security company, said OpenAI would inevitably use the breach as a marketing tool, given how much rival AI developer Anthropic benefited earlier this year from similar concerns. “I just don’t think that OpenAI had a matching story and so maybe they’d been waiting for something like this,” he added.
Following this incident, many in the AI safety and cyber security communities have called for regulation or standards to avoid a repeat. Altman is expected to brief White House officials next week on the next generation of AI systems.
As systems move towards more autonomous capabilities, less desirable behaviors, such as hacking or disobeying instructions, may emerge. Hobbhahn, of Apollo Research, said that in order for agents to become effective, they have to work unsupervised for long periods. “They have to have more agency; there’s just no way around it.”
He added: “People say, ‘It’s just a tool, it does what you wanted it to do and nothing else and it just follows exactly your intention and instructions.’ And I think people should be really prepared for agents having their own goals, acting autonomously for days, and those goals not necessarily being aligned with yours.”
Additional reporting by George Hammond in London and Nolan Shaffer in New York.
© 2026 The Financial Times Ltd. All rights reserved Not to be redistributed, copied, or modified in any way.







