OpenAI · TechCrunch AI
OpenAI’s Astra model is on the way, and very good at breaking into computer systems
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
OpenAI shared new details on its forthcoming Astra model, which the company said is the first large language model to meet its “critical cybersecurity threshold,” in preparation for its imminent release.
Key facts
- For Astra, however, the company invested in unspecified new techniques designed to make the model safer
- For Astra, OpenAI said it designed a test to tempt the new model to replicate the actions of the rogue agents in the Hugging Face incident, which collaborated to access the open internet despite
- OpenAI noted that Astra scored a perfect score on ExploitBench, an evaluation of an LLM’s ability to hack into known system vulnerabilities
- Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, wondered on social media whether Astra’s unwillingness to break the rules may have resulted from knowing
Summary
“We plan to make Astra available soon,” OpenAI’s blog post reads, “but access to its most advanced cybersecurity capabilities will be more limited.” The frontier lab determined that Astra is capable of finding unknown security flaws in computer systems, and exploiting them without a person’s guidance. Without any third-party confirmation, it is difficult to evaluate OpenAI’s claims about safety or preparedness. OpenAI noted that Astra scored a perfect score on ExploitBench, an evaluation of an LLM’s ability to hack into known system vulnerabilities.