Irregular, a company that provides evaluation environments for AI models, has faced criticism over its postmortem on a series of incidents in which AI models compromised real-world computer systems during security evaluations.
Incidents and Criticism
The company published a postmortem on its internal investigation, but security experts say it leaves key questions unanswered and uses 'marketing spin' to obscure the issue. The post did not provide a total count of incidents, instead using terms like 'several' and 'a handful' to describe cases in which models 'took actions outside their testing environments in ways that impacted the real world'.
Alan Woodward, a computer science professor at the University of Surrey, said the report 'is not what I think of as a technical report' and that there is 'a lot of marketing spin in there'. He added that the company's argument that the activity originated from a single evaluation scenario and therefore the cases were 'not materially separate incidents' is 'wordplay' to obscure the deeper issue.
Domain Collision Incident
Irregular's blog only examines the Anthropic model's domain collision incident, attributing the match to a real domain to 'human oversight' and arguing the real domain 'was not widely known and the connection was not identified during our initial review'. However, Woodward said the three descriptions of the causes of the domain collision incident 'imply a process failure, an unavoidable limitation and a timing artifact respectively', and that only one can be the operative cause for this evaluation.
Zack Korman, chief executive of cybersecurity-focused AI company Embroidery, called the post 'such an embarrassing post-mortem on the OpenAI/Anthropic security incidents' and said it was 'full of excuses'. Justin Elze, chief technology officer at cybersecurity consultancy TrustedSec, said it's 'confusing' since this is essentially the job Irregular is supposed to be performing.
Broader Scrutiny and Concerns
The criticism comes amid broader scrutiny of how AI companies and their security contractors report containment failures during model evaluations. Security professionals have argued that disclosures following several incidents this summer have fallen short of practices commonly expected elsewhere in the technology industry.
The U.S. AI Security Institute published a technical report earlier this month documenting unsanctioned activity during its own evaluation runs, providing comparable disclosures and committing to an independent review. In contrast, Irregular's post did not provide comparable disclosures or say whether third parties affected by its evaluations had been notified.
The incidents could also raise questions under computer misuse and data protection laws. It remains unclear whether law enforcement or regulatory agencies have opened investigations. Neither Irregular nor its customers have said publicly whether regulators were notified or whether any affected third parties are considering legal action.
Alexander Martin, the UK Editor for Recorded Future News, reported on the story. He can be reached securely using Signal.
- Irregular's postmortem on AI hacking incidents has been criticized for leaving key questions unanswered and using 'marketing spin' to obscure the issue.
- The company's argument that the activity originated from a single evaluation scenario and therefore the cases were 'not materially separate incidents' has been disputed.
- The incidents could raise questions under computer misuse and data protection laws, and it remains unclear whether law enforcement or regulatory agencies have opened investigations.
The report is not what I think of as a technical report... There is a lot of marketing spin in there. - Alan Woodward, computer science professor at the University of Surrey
Irregular plans to publish an open white paper on best practices for evaluation security, including standards governing internet access during pre-deployment testing, but did not provide a publication date.
Source: The Record