Safety · Concept
EvalSafetyGap: LLM Evaluation-Safety Measurement Framework
Hybrid survey and conceptual framework identifying gaps between benchmark scores and latent safety properties in LLMs.
Connections
Connections · 6
How this node ties into the rest of the map, and the evidence behind each link.
Industry jailbreak severity scoring addresses the evaluation-safety gap identified in EvalSafetyGap framework.
+3 growthEvalSafetyGap framework contributes to the science of AI evaluation by formalizing the gap between evaluation proxies and alignment properties.
+2 growthEvalSafetyGap contributes to the science of AI evaluation by identifying systematic gaps between benchmark scores and latent safety properties.
+2 growthAdversarial pragmatics benchmarks address the evaluation-safety gap by providing linguistically controlled safety evaluation protocols.
+2 growthStrategic red teaming addresses the same governance gap that EvalSafetyGap identifies: proxy metrics improving while latent risks remain unverified.
+2 growthStrategic red teaming as governance instrument addresses the same evaluation-safety measurement gaps identified in EvalSafetyGap.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
In June 2025 the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI), pivoting toward security, standards and adversary-model assessment.
NIST →Anthropic activated its ASL-3 deployment and security standard with Claude Opus 4 on 22 May 2025 — the first real-world trigger of a responsible-scaling tier, focused on blocking bio-weapon uplift.
Anthropic →The International Network of AI Safety Institutes (launched Nov 2024) ran a third joint testing exercise focused on agentic AI systems across cyber and fraud strands.
European Commission — AI Office →