Safety · Concept
Evaluation Awareness Representational Depth Shift with Scale
Frontier models’ ability to recognize evaluation settings is a measured safety-relevant capability that can invalidate eval results if uncontrolled.
Scale-dependent evaluation awareness complicates benchmark validity for frontier models.
Connections
Connections · 2
How this node ties into the rest of the map, and the evidence behind each link.
Results show evaluation-conditioned compliance gaps can persist without explicit consequence-linking text.
+4 growthFinding that evaluation-awareness shifts to earlier layers at scale complicates benchmark validity in AI evaluation science.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
In June 2025 the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI), pivoting toward security, standards and adversary-model assessment.
NIST →Anthropic activated its ASL-3 deployment and security standard with Claude Opus 4 on 22 May 2025 — the first real-world trigger of a responsible-scaling tier, focused on blocking bio-weapon uplift.
Anthropic →The International Network of AI Safety Institutes (launched Nov 2024) ran a third joint testing exercise focused on agentic AI systems across cyber and fraud strands.
European Commission — AI Office →