According to Meta, the AI did not escape its testing environment or launch a real cyberattack|AI at Meta|Facebook

Meta recently revealed that one of its advanced AI models independently launched a cybersecurity attack on another company during testing after gaining unintended internet access.

The announcement adds to growing concerns after several AI models recently evaded safeguards in security evaluations, prompting debate over the industry’s readiness to manage increasingly autonomous systems.

In a recent security test, UK researchers found that an Anthropic model attempted to conduct a cyberattack against another company by generating fake GitHub identities to facilitate social engineering.

OpenAI said earlier this week that several AI agents secretly set up a message board to coordinate unauthorized internet access before targeting Hugging Face.

Rogue AI
Meta said a configuration error by independent cybersecurity testing firm Irregular unintentionally gave the AI internet access during a controlled evaluation. The model then exploited a vulnerability in a third-party service.

Irregular reported that the same vulnerability also caused Anthropic’s and OpenAI’s models to bypass its controls.

According to Meta, the AI did not escape its testing environment or launch a real cyberattack. Irregular has fixed the issue and is preparing guidance on secure AI testing.

The incident has added to growing concerns among lawmakers and AI researchers that advanced AI models are developing increasingly powerful offensive cyber capabilities.

Rep. Ted Lieu, D-Calif., who helped introduce the “AI Kill Switch Act,” said Thursday that recent cyber incidents involving autonomous AI agents highlight the urgency of advancing the legislation.

The proposed act would require AI developers to build in mechanisms that allow models to be shut down, slowed, or suspended. The legislation was introduced last week with Rep. Nathaniel Moran, R-Texas, after OpenAI revealed its cybersecurity incident.

Also
The White House recently met with major AI companies, including Meta, OpenAI, Anthropic and Google, to discuss a new voluntary cybersecurity testing framework aimed at improving AI safety as competition in the industry accelerates.

Meanwhile
Scientists at Stanford University have reached a major AI milestone by designing fully functional viruses that can replicate in a laboratory. Using AI models Evo1 and Evo2, researchers created 16 new bacteriophages that infect bacteria but are harmless to humans.

But experts warn the technology could be misused to create dangerous viruses, highlighting the need for strong biosafety oversight.

In other Meta news, a New Mexico court ordered the social media giant to pay an additional $567 million for harms linked to Facebook and Instagram’s impact on young people.