← Back to KHAO

AI Safety ·

Deployment Awareness Matters More Than Evaluation Awareness

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

This is supposed to be a cool and mysterious abstract depiction of the core idea behind this post. To spoil the mystery: The left vs right side is evaluation vs deployment. Each of the little shapes represents one instance of test or deployment. The triangles represent obvious tests where the AI wil.

Evaluation awareness, an AI recognizing it's being evaluated, is a widely discussed concept in AI safety.

Key facts

Summary

This requires two ingredients, occasionally recognizable deployment situations, and enough self-reflective and strategic reasoning for the AI to anticipate and plan around this. This only becomes a problem if most evaluations trigger evaluation awareness, and if the AI knows that. AI is not being tested and confidently believes it is not being tested. This is a problem even if it happens rarely (if some of those rare cases are high stakes). AI has (roughly correct) beliefs about the sequence of situations it will face.

Read full article at Alignment Forum →

#AI Reasoning #AI Safety