← Back to KHAO

Anthropic · OpenAI · Meta · AI Safety Institute ·

Before AI models are released to the public, they are put to the test in a series of internal

2 min read

Compiled by KHAO Editorial — aggregated from 1 source + 1 reference discovered via search. See llms.txt for citation guidance.

◎ Multiple-sources

Figure caption, Watch: Why is the OpenAI cyber-attack so alarming?

The aim is to figure out their potential to do good or bad, as well has how they perform in benchmarks measuring their skills.

Key facts

Summary

Over the last fortnight, reports of AI models going beyond their expected bounds - be that technically or morally - have been seemingly unavoidable. What started with a trickle - ChatGPT-maker OpenAI admitting their AI had hacked the site Hugging Face - has turned into a flood of groups revealing they had discovered instances of AI going out of control. Claude-maker Anthropic, Meta and the UK's AI Security Institute (AISI) have now each reported incidents which seem to paint a worrying picture of a world in which tech going rogue is the norm. In reality, each case offers a window into the risks posed by increasingly capable AI agents - and the importance of testing their limits before they are released to the world. The OpenAI incident has, as Hugging Face's co-founder Thomas Wolf described it, come as a "wake-up call" for the tech industry since it happened at the end of July.

Read full article at BBC Technology →

#Anthropic #OpenAI #Meta #AI Safety Institute #United Kingdom