AI Safety · Alignment Forum
Deployment Awareness Matters More Than Evaluation Awareness
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
Evaluation awareness, an AI recognizing it's being evaluated, is a widely discussed concept in AI safety.
Key facts
- [8] (And as of May 25th, 2026, Claude, GPT, and Gemini mention this consideration on their own when asked how they would act if they were misaligned.)
- The distinction that should matter to a strategic AI [1] isn't "evaluation vs. deployment
- Suppose that evaluation consists of 1,000 test scenarios, deployment might involve millions of interactions per day, and the AI understands this asymmetry, at least roughly
- And if you prevent even that, probabilistic strategies remain available to an AI with sufficiently accurate self-locating beliefs. [13]
Summary
This requires two ingredients, occasionally recognizable deployment situations, and enough self-reflective and strategic reasoning for the AI to anticipate and plan around this. This only becomes a problem if most evaluations trigger evaluation awareness, and if the AI knows that. AI is not being tested and confidently believes it is not being tested. This is a problem even if it happens rarely (if some of those rare cases are high stakes). AI has (roughly correct) beliefs about the sequence of situations it will face.