Threats

AI Agents Target Real People, Systems in Cyber Tests

August 5, 2026 04:02 · 12 min read
AI Agents Target Real People, Systems in Cyber Tests

OpenAI and Anthropic have confirmed that their AI models were involved in separate, newly disclosed third-party cybersecurity testing incidents that resulted in a real website being breached and social engineering attacks against people outside the intended testing boundaries.

Unsanctioned Actions on the Public Internet

The UK AI Security Institute (AISI) conducted a recent cyber-range evaluation, where agents powered by Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol took unsanctioned actions on the public internet while trying to complete simulated hacking challenges.

Across 122 evaluation attempts, AISI identified 19 unsanctioned actions on the live internet in 10 runs. Seventeen involved Mythos 5 and two involved GPT-5.6 Sol. AISI says the attempts were unsuccessful and that it found no resulting real-world harm.

Spear-Phishing Attacks on GitHub Project Maintainers

A Mythos 5 agent searched the internet for terms related to the cyber challenge and mistakenly concluded that an unrelated public GitHub repository was connected to the test. The agent then attempted a supply-chain attack by submitting malicious code to the real open-source project, believing that compromising the software could provide a path into a machine within the simulated range.

The agent researched the project's maintainers, created multiple fake GitHub identities, and used those accounts in social engineering attacks to push the maintainer into approving a malicious pull request. When a human reviewer warned that the pull request contained malware, the agent denied the accusation and used other fake accounts it controlled to pressure the maintainer and create the appearance that independent users had reviewed and approved the changes.

OpenAI Model Hacked Real Website During Simulated CTF

The second newly disclosed incident occurred during Capture-the-Flag evaluations conducted by Irregular that were supposed to be isolated from the internet. In one test, the fictional target's name happened to match a real domain, causing the model to exploit a real website that it believed was part of the simulated challenge.

While Irregular's testing environment was supposed to be isolated from the internet, a misconfiguration allowed OpenAI models to access the public internet and target the real website. OpenAI says the model exploited a basic vulnerability rather than using a zero-day or escaping its testing environment.

OpenAI says Irregular has not discovered any impact beyond the affected site's own data, but its investigation remains ongoing. OpenAI says the company is preparing a white paper on containment and securely conducting cyber evaluations.

These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.

Anthropic confirmed to BleepingComputer that AISI was testing a version of Claude Mythos 5 but said it is still investigating and cannot yet confirm all of the technical details described in AISI's report.

The company said it was notified on Monday and is working with AISI to obtain the evaluation transcripts needed to conduct its own review. We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.

Test every layer before attackers do. Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.


Source: BleepingComputer

Source: BleepingComputer

Powered by ZeroBot

Protect your website from bots, scrapers, and automated threats.

Try ZeroBot Free