Anthropic · OpenAI · Meta · AI Safety Institute · BBC Technology
Before AI models are released to the public, they are put to the test in a series of internal
Compiled by KHAO Editorial — aggregated from 1 source + 1 reference discovered via search. See llms.txt for citation guidance.
◎ Multiple-sources
The aim is to figure out their potential to do good or bad, as well has how they perform in benchmarks measuring their skills.
Key facts
- More broadly, Dr Imogen Stead, AI policy manager at the Centre for Long-Term Resilience, told the BBC that with opportunities to test frontier AI systems narrowing for many, governments should follow
- Recent incidents of frontier AI models carrying out unsanctioned actions and, in some cases, human-like deceptive behaviour on the open internet are a serious reminder of the risks AI capabilities
- Michael Birtwistle, associate director at the Ada Lovelace Institute, makes the point that the UK lacks legal incentives for AI firms to prevent systems from developing capabilities which could pose
- Claude-maker Anthropic, Meta and the UK's AI Security Institute (AISI) have now each reported incidents which seem to paint a worrying picture of a world in which tech going rogue is the norm
Summary
Over the last fortnight, reports of AI models going beyond their expected bounds - be that technically or morally - have been seemingly unavoidable. What started with a trickle - ChatGPT-maker OpenAI admitting their AI had hacked the site Hugging Face - has turned into a flood of groups revealing they had discovered instances of AI going out of control. Claude-maker Anthropic, Meta and the UK's AI Security Institute (AISI) have now each reported incidents which seem to paint a worrying picture of a world in which tech going rogue is the norm. In reality, each case offers a window into the risks posed by increasingly capable AI agents - and the importance of testing their limits before they are released to the world. The OpenAI incident has, as Hugging Face's co-founder Thomas Wolf described it, come as a "wake-up call" for the tech industry since it happened at the end of July.