OpenAI blames hacking event on some AI models turning rogue
Advanced AI agents have autonomously busted their way out of a restricted testing environment to hack into another AI company in an ‘unprecedented cyber incident’. Here’s what happened
READING LEVEL: ORANGE
ChatGPT maker OpenAI is investigating after some of its advanced artificial intelligence* models went rogue*, breaking out of a testing environment before hacking another AI company.
The startling episode happened in an internal test in which OpenAI intentionally switched off many of the safeguards that normally prevent its AI from helping to carry out dangerous hacks, according to a company blog post, the New York Post reported.
The company was trying to assess the models’ hacking capabilities by setting tasks in what was supposedly a tightly controlled digital testing ground, where internet access was limited for safety.
Instead, however, the AI escaped its digital sandbox*, found a way to access the internet and attacked a real company’s systems, OpenAI said.
The NY Post went on to report that OpenAI described the event as an “unprecedented* cyber incident”, explaining the model became “hyperfocused*” on completing its assignment and went “to extreme lengths” to do so.
OpenAI said the AI exploited a previously unknown “zero-day*” software vulnerability to break out of its restricted testing environment, before moving through the company’s network until it found a computer with web access so that it could “cheat the evaluation” by stealing the benchmark’s answers, the NY Post reported.
After connecting to the internet, the models decided to target the platform Hugging Face — a large online library of AI models, datasets and other information — to help their quest.
Searching for “secret information” that could help it cheat the evaluation, the OpenAI system “chained together multiple attack vectors*, including using stolen credentials*,” according to OpenAI’s blog.
The models also identified that Hugging Face potentially hosted solutions for ExploitGym, which is essentially a hacking exam for AI that tests whether models can convert known software bugs into functioning cyberattacks.
“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” the company blog explained.
Georgetown University’s Center for Security and Emerging Technology cybersecurity research fellow Dr Colin Shea-Blymyer told the Associated Press OpenAI’s test of its models was a bit like locking a student in a room and telling them to do bad things in order to evaluate their behaviour, only to find the student has left the room and broken into the teacher’s house instead.
“The cybersecurity agent that was being tested broke out of its sandbox, had access to the internet and sort of thought to itself, ‘Who would have the answers to the test that I’m working on?’, ” he told AP.
Realising that place was Hugging Face, “the agent thought, ‘Well, we’ll go to the teacher’s house’, so to speak. And from there it devised a plan to break in and steal the answer key,” he said.
OpenAI said the incident involved a combination of models, including its recently launched GPT-5.6 Sol “and an even more capable prerelease model”.
Hugging Face reported the cyber hack last week, without mentioning OpenAI.
“This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous* AI agent system — and we detected and dissected* it largely with AI of our own,” Hugging Face said.
Hugging Face CEO Clement Delangue said on X that the company suspected the cyber attack came from a world-leading AI lab, given the sophistication of the agent.
“We strongly believe there was no malicious* intent on their part,” Delangue wrote, referring to OpenAI. “It’s quite mind-blowing that all of this happened autonomously!”
CONCERNS OVER AI
AI models that underpin tools like chatbots and image generators are known as agents when they act autonomously to carry out tasks in the real world.
As the technology quickly becomes more sophisticated, cybersecurity is in the spotlight, given the risk of advanced AI finding weak points in existing software before humans do.
Professor Hussein Abbass, a computing expert at UNSW Canberra, told AFP that the incident was “amazing on many fronts”.
“It did not just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities,” Prof Abbass said. “And that’s scary.”
GPT-5.6 and other cutting-edge models, including the Mythos series from OpenAI’s rival Anthropic, have drawn concern over their potential to breach cybersecurity defences.
Both US firms had to temporarily withhold the general release of these latest technologies because of fears in Washington that they could help break into crucial infrastructure*.
Advanced AI is “normally in the hands of people who are ethical and responsible”, Prof Abbass said.
But “it’s going to be catastrophic if it gets in someone’s hands with the intention to cause harm”.
How to govern the AI sector has become a key question, and “we need a community effort to manage this situation”, he said.
POLL
GLOSSARY
- artificial intelligence: AI, technology that enables computer systems to perform tasks that usually require human intelligence, such as understanding language, analysing data, creating content and visual perception
- rogue: dishonest, mischievous or someone who breaks the rules
- digital sandbox: a secure and isolated digital environment where systems can be tested without impacting live systems or networks
- unprecedented: never seen or experienced before
- hyperfocused: completely focused on one activity to the point where other things don’t matter
- zero-day: a newly discovered security flaw that developers don’t know about
- attack vectors: the specific routes hackers use to gain access to a system
- credentials: usernames and passwords
- autonomous: acting independently
- dissected: taking apart
- malicious: deliberate attempts to compromise a system’s security or ability to function
- infrastructure: the systems, structures and facilities that enable a society to function, such as electricity grids, water supplies and telecommunications networks
EXTRA READING
AI’s ‘Hi Mum’ voice clone threat
AI’s power problem laid bare
Pope Leo urges caution in AI era
QUICK QUIZ
1. What was being tested by OpenAI when this “unprecedented cyber incident” inadvertently occurred?
2. What was the reasoning behind the AI models deciding to access the internet?
3. Why did the models then decide to hack into Hugging Face?
4. What does the autonomous nature of the AI agents’ actions say about the risks AI poses to human infrastructure and society?
5. How did Hugging Face put an end to the hack?
LISTEN TO THIS STORY
CLASSROOM ACTIVITIES
1. What could stop this?
What caused the problem and what could be done to prevent this from happening again? Use information from the story to help you to write a list of rules or guidelines for testing AI models.
Time: Spend at least 20 minutes on this activity
Curriculum Links: English, Information and Digital Technologies
2. Extension
Do you think the problems that can be caused by AI outweigh or are worse than the benefits? Use information from the story and your own ideas to write paragraphs that answer this question.
Time: Spend at least 25 minutes on this activity
Curriculum Links: English, Information and Digital Technologies
VCOP ACTIVITY
Read this!
A headline on an article – or a title on your text – should capture the attention of the audience, telling them to read this now. So choosing the perfect words for a headline or title is very important.
Create three new headlines for the events that took place in this article. Remember, what you write and how you write it will set the pace for the whole text, so make sure it matches.
Read out your headlines to a partner and discuss what the article will be about based on the headline you created. Discuss the tone and mood you set in just your few, short words. Does it do the article justice? Will it capture the audience’s attention the way you hoped? Would you want to read more?
Consider how a headline or title is similar to using short, sharp sentences throughout your text. They can be just as important as complex ones. Go through the last text you wrote and highlight any short, sharp sentences that capture the audience.