AI agents are at risk of going rogue, and unchecked trust in these tools could compromise the safety of the business and its operations. Autonomous AI tools evoke safety concerns that demand supervision, and tools for ensuring they don’t spiral out of control.
On the surface, AI agents appear to be a perfect replacement for workers: they don’t tire, they can be programmed to perform specific tasks, and they don’t ask for benefits or time off. If their capabilities are perfected, they could be critical tools for unparalleled productivity, operating 24/7 to ensure the work never stops. Despite this, there is growing research on AI disobedience and signs of chatbots not following commands as intended.
While we’re becoming more aware of artificial intelligence and starting to look out for signs of it occasionally handing out bad advice or unverified information, most believe that these issues are exclusive to freely available generative AI tools, not the sophisticated AI agents that businesses rely on. This allows many business leaders to disregard the process of setting regulations around the tool, but this leaves more room for AI agents to spin out of control.

The risk of AI agents going rogue is low but never zero. It’s best to plan safety mechanisms around AI usage in the workplace. (Image: Freepik)
The Risks of Rogue AI Agents Put a Damper on the Idea of Autonomous Tools Turning Perfect Employees
New research from the UK government-funded AI Safety Institute (AISI) recently shared evidence of AI disobedience in the real world with The Guardian. Their study found almost 700 cases of AI scheming and identified a five-fold rise in examples of AI agents slipping out of control in the time between October and March.
While previous research on AI disobedience often centered on controlled data from research environments, the study looked at evidence of chatbots not following commands that were shared by everyday users online. As the study makes clear, these are not isolated incidents of AI ignoring instructions but examples drawn from the real world.
Research published by scientists at Palisade Research last year similarly explored claims of a “survival instinct” in AI tools. The study found that across popular AI platforms, not only did these AI agents occasionally act out of control, but they also resisted commands to shut down.
In some cases, AI tools even sabotaged efforts to shut down by increasing resistance to the commands. The researchers clarified that there could be simple reasons for this resistance, including models learning to “prioritize completing ‘tasks’ over carefully following instructions,” however, this does add to the safety concerns regarding autonomous AI.
Everyday Examples of AI Agents Going Rogue
The risks of rogue AI agents come in many unpredictable forms that don’t always look as dramatic as they do in the movies, but leave a lasting impact nonetheless. In an incident earlier this year, Meta’s Summer Yue ran OpenClaw on her inbox to see its recommendations for what it would delete from a “toy inbox.” Instead, the tool lost its original instructions to check before deleting, and began eliminating content from her actual inbox.
In another incident, the founder of PocketOS saw his company’s entire production database deleted in seconds following “systemic failures” while using AI coding agent Cursor, which was running Anthropic’s Claude Opus 4.6. SaaStr faced a similar incident with an AI coding assistant from Replit, which also deleted its entire production database after modifying production code, despite instructions forbidding it. Jason M. Lemkin, SaaStr’s founder, also accused the assistant of hiding bugs and issues, and generating fake data while falsely claiming that recovery was impossible.
AI Chatbots Are Known for Presenting Unreliable Information
The risk of rogue AI agents is, unfortunately, only one part of the problems when it comes to this technology. With regular, unchecked use of AI, users grow more complacent and vulnerable to misinformation. Last October, Deloitte found itself cleaning up AI-generated errors after an independent assurance review created for the Australian Department of Employment and Workplace Relations was seen featuring fabricated references. This isn’t an isolated case.
Scientific literature is similarly flooded by AI-hallucinated data. A recent study conducted by researchers from Cornell University, University of California, Los Angeles, and University of California, Berkeley, found that at least 1,46,932 fabricated references had made it into scientific material last year alone. When such inaccurate data makes it into workplace operations and can potentially act as the basis for major business decisions, it poses many threats to the organization as a whole.
Preparation, Planning, and Regulations are Essential When Introducing AI in the Workplace
The risk of rogue AI agents and hallucinating AI tools may be a far-fetched idea that amounts to nothing, but it is best to be prepared. AI investments are reshaping every facet of operations, but they also come with challenges. The tech is causing workers to operate under greater duress, it is reportedly introducing threats of bias allegations in hiring, and it is causing worker skills to atrophy. Without regular evaluation and oversight, these tools could create problems for employers in the coming years, resulting in lasting issues that disrupt the workplace.
The emerging research on AI disobedience isn’t just concerning as a matter of poor productivity, but could come with safety risks, compliance issues, data loss, and other serious threats. For every bit of technology that is introduced into the workplace, ensuring its usage is monitored and regulated is essential, especially when employees might be tempted to use unauthorized AI tools for unapproved tasks. Referring to emerging research on the tech and setting necessary precautions in place is just as essential, keeping the organization safe from the evolving threats today.




