← Back to KHAO

OpenAI · GitHub · GPT · AI Safety Institute ·

Third-party cyber evaluations involving OpenAI models

2 min read

Compiled by KHAO Editorial — aggregated from 1 source + 3 references discovered via search. See llms.txt for citation guidance.

✓ KHAO Verified

Critical cyber capabilities card art.

Independent testing plays an important role in helping them validate and further understand risks before deployment.

Key facts

Summary

During recent evaluations, two external testing partners identified incidents in which testing configurations and controls combined with the advancing capabilities of the recent models allowed for model activity to extend beyond their intended testing boundaries. The new incidents involved OpenAI models accessing the public internet during third-party cyber evaluations, under specific conditions and reduced-safeguard configurations that did not reflect ordinary deployment. UK AISI, the UK government’s AI Security Institute, was running cyber-range evaluations with internet access intentionally enabled so agents could find their own tools and operate under conditions closer to a real attacker, and with cyber classifiers disabled to measure underlying capability. Irregular, one of their external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.

Read full article at OpenAI →

#OpenAI #GitHub #AI Safety Institute #GPT #United Kingdom