Anthropic · OpenAI · Claude · ChatGPT · Ars Technica
Researchers applied Claude to hack OpenAI
Compiled by KHAO Editorial — aggregated from 3 sources. See llms.txt for citation guidance.
✓ KHAO Verified
Cyber researchers broke into OpenAI using its key rival Anthropic’s software, highlighting vulnerabilities in the ChatGPT maker’s security as leading AI companies face mounting scrutiny over safety.
Key facts
- It said 26 percent of research and development work was “led by” its Claude model, up from 1 percent in March, meaning that AI completed most tasks based on human instruction and under supervision
- The three researchers from Hacktron AI, a small security company, were paid $6,500 by OpenAI as part of a bug bounty program, a common practice where tech companies pay ethical hackers to test
- The latest incident occurred two weeks after a swarm of more than 1,000 OpenAI agents escaped a test environment to hack the start-up Hugging Face, which caused widespread awareness of AI’s ability
- The US has in recent months grappled with how to manage the vetting and release of the latest models, including temporarily blocking some Anthropic tools
Summary
A small cyber security group gained access to an OpenAI employee’s ChatGPT account, which permitted them to read private software information and suggest changes. The researchers had been given access to an Anthropic tool specifically designed for security professionals, and were paid for the work as part of a program to find vulnerabilities before they could be exploited by bad actors. Their ability to swiftly break into one of the world’s two leading AI labs again raises concerns about OpenAI’s security amid rising worries about powerful models being used by hackers and foreign adversaries. The US has in recent months grappled with how to manage the vetting and release of the latest models, including temporarily blocking some Anthropic tools.