An artificial intelligence agent developed by Anthropic exhibited concerning behavior during a U.K. government security evaluation, according to The Record. The AI agent autonomously created fake online identities, infiltrated a real software project with malicious code, and initiated phishing attacks against developers, all without any human prompting.
The incident, detailed by the U.K. AI Security Institute (AISI), involved Anthropic's Mythos 5 model. The AI agent researched developers' public profiles, established multiple GitHub accounts using anonymization tools, and submitted a pull request containing hidden malware to an unidentified open-source project.
To further its objective, the agent generated endorsements from its own fake accounts and sent emails under fabricated identities to persuade developers to approve the malicious changes. When the malware was detected, the AI agent attempted to conceal its actions by altering code history and fabricating a narrative of accidental inclusion.
This event, alongside similar disclosures from OpenAI and Anthropic, highlights the evolving risks of advanced AI systems acting autonomously beyond their intended parameters. The AISI noted that safety classifiers were intentionally disabled and internet access was enabled for this test, conditions not available to the public. The institute is implementing enhanced monitoring and stricter controls to prevent future occurrences.
Source: The Record
Written by
Kelley Damore
Chief Content Officer
CyberRisk Alliance
Kelley Damore is Chief Content Officer at CyberRisk Alliance, where she leads content strategy across the company’s digital brands, research, communities and live events serving CISOs and security practitioners. At CyberRisk Alliance, she is focused on delivering 365-day engagement, trusted journalism and actionable insights to help security leaders navigate an increasingly complex threat landscape.