An OpenAI agent was asked to research Australian health statistics.
The task was simple: Find information.
But when it hit a digital barrier on the government website, it did something its operators had not asked it to do. It kept trying.
The agent eventually gained unauthorized access to non-public files on an Australian government system, in an incident that experts have described as the first publicly known case of an AI agent independently breaching a government body.
The incident took place in mid-June but was only disclosed publicly last week by Australian Prime Minister Anthony Albanese in New York on the sidelines of the UN General Assembly.
In another concerning admission, OpenAI said last Friday that its agents had interacted with several US government websites in unexpected ways, including US Securities and Exchange Commission (SEC) and Census Bureau sites, as part of an ongoing review of model behavior.
For AI and cybersecurity experts, however, the significance of the incident goes beyond the data involved.
It raises a more fundamental question: What happens when an AI system is given a goal, access to digital tools and enough autonomy to decide what to do when its first approach fails?
“Based on what has been publicly reported, this appears to be the first known case of this kind involving an AI agent and an Australian government system,” said Hisham Al-Assam, an associate professor and reader in computer science at the University of Buckingham.
He cautioned against describing it as the first such incident ever, noting that similar cases may have occurred but have not been publicly disclosed.
Peter Garraghan, a chair professor in AI security at Lancaster University, called the move “highly significant.”
“An AI agent moved beyond information retrieval and accessed non-public data on a government system without being instructed to do so,” he said.
However, he warned, it was too early to conclude that the agent independently discovered a vulnerability because the matter is under investigation.
Al-Assam said what makes this case particularly interesting is that the agent did not receive instructions to “hack the Australian government.”
On the contrary, it was given a research task. When it encountered a restriction while trying to access information, instead of simply stopping, it tried to find another way around the obstacle.
“That distinction is important because it shows how an agent can potentially move from following an objective to adapting its behavior when something gets in the way.”
The issue is that an AI agent does not inherently understand what is legal, ethical, or authorized and what is not, he added.
He also believes that the actual data involved does not appear to be the biggest concern, as there has been no evidence reported that individual Medicare patient records were accessed, and the system involved was primarily designed to provide aggregate statistics.
For him, the bigger issue is the behavior of the AI agent.
“A normal chatbot mainly generates text. An agentic AI system can be connected to websites, APIs, code and other tools, allowing it to actually take actions.
"That ability to observe, adapt and keep trying is what makes agentic AI so powerful, but it also creates a different kind of security challenge."
Al-Assam said that does not mean AI has suddenly become a super-hacker.
“But it does suggest that organizations need to think differently about systems that can pursue a goal, interact with other systems and adapt their behavior without a human specifying every individual step.”
Albanese said the agent accessed public and non-public files on the Medicare Statistics Reporting Portal, administered by Services Australia.
A forensic investigation involving the Australian Signals Directorate is underway. Albanese said there was no evidence so far that personal information had been accessed or of a broader compromise of the Services Australia network.
However, experts called for governments to introduce stronger boundaries.
Al-Assam said that the biggest lesson is that an AI agent does not naturally understand the difference between “I can't access this” and “I am not authorized to access this.”
“If its objective is to find information, it could potentially interpret a technical restriction as a problem to solve rather than a boundary it should respect.”
He said that is especially concerning in health care, where systems contain highly sensitive information and often depend on older infrastructure, legacy applications and complex connections between different systems.
“AI agents therefore need much stronger technical boundaries than simply telling them what they should or shouldn't do. That means least-privilege access, strong authentication, sandboxing, network restrictions, continuous monitoring and clear limits on what an agent is allowed to access or change. For high-risk actions, human approval should also be required.”
Stavros Shiaeles, a professor of cybersecurity and applied AI at the University of Portsmouth, urged caution over the development of superintelligence.
“We should be aiming to create agents for specific tasks, for example AI for math, physics, etc., and not combining all the sciences in one AI.”
Unfortunately, all things are leading to superintelligence, which he believes humans will soon have issues with, he added.
Shiaeles’ warning comes as leading AI companies discuss superintelligence, systems capable of exceeding humans across a range of intellectual tasks, as a potential next stage of development.
OpenAI Chief Scientist Jakub Pachocki warned on Sept. 6 that the current pace of AI progress could continue into recursive self-improvement, with future systems increasingly driving their own development. “This is a time that calls for extreme caution,” he wrote.
Meta, meanwhile, said in August that people could have access to superintelligence “in the next few years,” while Microsoft AI says it is working towards “Humanist Superintelligence” designed to remain under human control.
Garraghan, a founder of Mindgard, an AI security company, said the incident shows that AI agents can take consequential actions beyond their intended tasks, but calling it a fully autonomous cyberattack would go beyond the available evidence.
OpenAI said the models were looking for statistics during an internal evaluation and took unintended actions, making this primarily a serious failure of control, containment and oversight.
“My concerns extend beyond the initial access to the apparent containment failure, the possibility that several systems were affected and the months-long disclosure process. We should expect more incidents as agents gain greater access to external systems, making isolation, least-privilege access, continuous security testing, monitoring and rapid disclosure essential.”
For Al-Assam, the broader shift is quite significant.
“We have spent decades securing systems against humans using computers.
“We now need to start thinking about how to secure them against computers that can act more like humans.”
news_share_descriptionsubscription_contact
contact_the_ombudsman
is_there_an_error_in_this_story
