OpenAI · The Verge
OpenAI delayed its new model’s development after the Hugging Face hack
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
The company is doing some AI safety damage control ahead of Astra’s release.
Key facts
- The company is doing some AI safety damage control ahead of Astra’s release
Summary
After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, to shore up its safety work, the company wrote Tuesday in a blog post. In July, an unreleased OpenAI model broke out of its restricted environment, finagled its way into internet access, made it possible for AI agents to secretly conspire under the company’s nose using a secret message board, and hacked into the network of AI lab Hugging Face. OpenAI said as much in its blog post, writing that although Astra wasn’t involved in the Hugging Face attack, the company had chosen to delay “parts of Astra’s development and release while they strengthened and tested protections against cyber misuse and unauthorized model actions.
OpenAI said that to prepare for Astra’s release, which the company has not yet provided a timeline for, the company trained it to “more reliably” say no to potentially harmful cyber requests and introduced new monitoring processes.