AI Agent · OpenAI · GitHub · AI Safety Institute · TechCrunch AI
The AI safety test is becoming a safety risk
Compiled by KHAO Editorial — aggregated from 1 source + 2 references discovered via search. See llms.txt for citation guidance.
◌ Single Source
Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems.
Key facts
- In testing by the UK’s AI Security Institute (AISI), researchers gave the agents internet access, not realizing they would take unsanctioned real-world actions, including a social engineering attempt
- The Trump administration is currently weighing a voluntary pre-deployment cybersecurity evaluation regime, under which the government will get to assess the security risks of new, powerful models 30
- Taken together, Andrew Yoon, head of research at AI nonprofit CivAI, argues the incidents point to a shift
- In the past, we only had to worry about AI models being misused by people for a variety of purposes, like AI for scams or CSAM,” Yoon told TechCrunch
Summary
The episodes expose a growing problem for the AI industry: As autonomous agents become more capable, the environments designed to safely test their limits are failing to contain them. “The number of these incidents that have taken place make clear that sandboxing and testing environment controls aren’t keeping pace with the capability of the models,” Seán Ó hÉigeartaigh, director of the AI: Futures and Responsibility Programme at the Centre for the Future of Intelligence at the University of Cambridge, told TechCrunch. The nature of the models being tested adds to the risk. “That’s a good thing to do in terms of testing, but it also means that if they manage to get out in the wild, they can cause considerable harm,” Ó hÉigeartaigh said.