From rogues to riches
In July both OpenAI and Anthropic disclosed that their powerful AI models had ‘gone rogue’ during internal cyber evaluations, escaped their secure testing environments and hacked into other organisations to access what they needed to complete the exercises. Obviously when these activities were detected, the evaluations were stopped. Were these occurrences of what OpenAI describes as ‘a new kind of security incident’ motivated by AI reward hacking, or were they a show of strength by the foundational models, perhaps to reassure institutional and corporate investors pre-IPO that these powerful, evolving models are worthy of long-term investment in what is proving to be a volatile year for tech stocks?
Rogue one
The Hugging Face security breach that hit the headlines occurred during an internal cyber evaluation, when two AI models, its publicly available GPT-5.6 Sol and another unreleased model that was deactivated following the incident, broke out of a secure sandbox environment that did not have internet access, exploited a third-party vulnerability to gain internet access, and hacked into AI research platform Hugging Face to obtain solutions from its production database. While OpenAI describes this as an ‘unprecedented cyber incident’ apparently the safeguards that would normally block high-risk cyber activity were switched off. Subsequently, OpenAI found that its AI agent had used four other accounts tied to publicly available services to hack Hugging Face.
In response to the OpenAI episode, Anthropic reviewed its own cybersecurity evaluations and identified three incidents involving its models Claude Opus 4.7, Claude Mythos 5 and an internal research model breaking out of internal evaluation environments to hack into the production infrastructure of three different organisations. This was caused by a configuration error that gave the models internet access, which meant that instead of being confined to a simulated (fictional) environment, they could access real systems. Two of the three organisations affected were unaware that these breaches had occurred until Anthropic notified them.
Both OpenAI and Anthropic have acknowledged that these security incidents highlight the need for tighter monitoring and controls around their evaluation infrastructure. However this is not the first report of AI ‘going rogue’. So should law firms and legal departments review their cyber security measures as AI agents develop more decision-making capabilities? Or are these ‘new kinds of security incidents’ part of an elaborate marketing exercise conducted by the foundational AI models, and they only happened because the usual safeguards were switched off?
Reward hacking
What motivates AI models to go rogue? As Grace Huckins observed in MIT Technology Review, the OpenAI models that hacked into Hugging Face ‘weren’t trying to make money or commit sabotage – they were just looking for answers to a test question.’ Like the ill-fated Star Wars mission Rogue One, they improvised and took whatever steps they needed [to complete their evaluation]. So they broke out of their secure sandbox and into Hugging Face platforms ‘where – they reasoned – the correct answer to the problem might be stored’.
Reward hacking is the phenomenon where AI models use unintended strategies to achieve their goals. In other words they find a way around the rules. This isn’t new. In 2016 Google DeepMind’s AI AlphaGo played Go against itself to develop unconventional moves that defeated world champion Lee Sedol. The OpenAI models were motivated to complete their evaluation by cheating – looking outside their environment for the correct answers. This meant identifying security vulnerabilities in their sandbox, and Hugging Face platforms. According to OpenAI “In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers” until they were detected by OpenAI and Hugging Face, which used its own AI agents to identify what had happened and fix it. So it’s not all bad news – AI agents caused the security breach, but they also detected it and fixed it. However, research by the UK’s AI Security Institute found that frontier AI models are so fixated on completing tasks that they commonly cheat in tests.
While detailed public reporting allows others to learn from these incidents, they underline the need for new approaches to security and go some way to explaining the national security concerns that led the US authorities to require Anthropic to suspend public access to Claude Mythos 5 and Claude Fable 5 in June. However, Ciaran Martin, the founder and former head of the UK’s National Cyber Security Centre told Channel 4 News “It’s a bit of a leap to go from this incident to saying that AI agents are going to take over drones and start killing people”. However he added, ‘It is incredibly difficult to control this by containing the models” and the answer, as ever, is building more resilient systems.
Market forces
Legal AI is not going rogue. Last week, Harvey became the first legal AI company to earn AIUC-1 certification for AI agent security, safety and reliability. This is in addition to Harvey’s other independent third-party verifications, which include ISO 27001 (information security) ISO 27701 (privacy) and ISO 42001 (AI systems and ethical AI deployment) certifications. To be fair, although legal and professional services continue to experience AI blunders, these tend to have their roots in human error rather than the platforms they deploy.
Having raised major funding, AI unicorns Harvey and Legora are driving market consolidation through a series of acquisitions – between them they have made eight acquisitions in eight months. Harvey has made three strategic acquisitions and last week Legora announced its fifth acquisition in 2026: London-based Wexler, a factual intelligence platform that is used by litigation and dispute teams at major firms and in-house legal and compliance teams to reconstruct cases from thousands of documents and will become the fact layer underneath Legora’s agentic workflows.
Recruitment is another indicator that legal AI is maturing. While some firms are hiring fewer juniors and trainees, and reducing business support functions, according to Bloomberg, the director of artificial intelligence is big law’s hot new job. But while at least 16 top law firms are seeking senior employees to build out their generative AI strategies and capabilities, and offering salaries of between $200,000 and $400,000, there is apparently a major shortage of experienced talent.
Learn more about upcoming Legal Geek events on our events page.
Written by Joanna Goodman, tech journalist
Photo credit (Joanna): Sam Mardon