Safety · Concept
Science of AI Evaluation
Empirical evaluation practice spanning cyber capability gaps, persuasion experiments, and governance measurement validity.
Connections
Connections · 45
How this node ties into the rest of the map, and the evidence behind each link.
ECAISA proposes explicit epistemic norms, scoring, and disclosure practices for safety/alignment evaluation work.
+5 growthIRT supplies latent safety factors, efficient adaptive item selection, and sandbagging-oriented model audits.
+5 growthHigh ambiguity rates imply current automated safety suites are weak evidence for SLM safety claims.
+5 growthPluralis v0.1 extends AI safety evaluation to multicultural, multimodal, multilingual contexts beyond Western-centric defaults.
+5 growthAISI's Frontier AI Trends Report exemplifies evidence-based AI evaluation science applied to frontier model assessment.
+4 growthUnderstanding that LLM reasoning is pattern-matching rather than abstract world modeling informs how AI evaluation science should be designed.
+4 growthMedFailBench reframes medical evaluation around which safety boundary failed rather than only knowledge accuracy.
+4 growthIRT supplies psychometric structure and adaptive item selection for safety eval science.
+4 growthThe Inattentional Gap finding decouples benchmark safety from real-world safety, demanding new evaluation science that tests for unspecified hazards.
+4 growthReasoning consistency scanning contributes to the science of AI evaluation by providing a tractable method for auditing chain-of-thought validity.
+4 growthThe methodology treats decision-maker understanding as an assessable safety object alongside system cards and safety cases.
+4 growthAda Lovelace Institute published commentary on strengthening the science of AI evaluation to bring clarity to AI risks and benefits.
+3 growthSystems-safety methods applied to agentic AI strengthen the science of AI evaluation by surfacing risks missed by model-level testing.
+3 growthText-based AI Safety Gridworlds provide controlled evaluation infrastructure for studying reward hacking, strengthening the science of AI evaluation.
+3 growthNRT-Bench advances the science of AI evaluation by providing objective harm signals rather than LLM-judged text for safety-critical agent assessment.
+3 growthIndustry-wide jailbreak severity scoring standardizes evaluation of adversarial attacks across AI systems.
+3 growthGemma 4 establishes new performance benchmarks across STEM, multimodal, and long-context tasks, advancing evaluation science.
+3 growthECAISA proposes epistemic norms and independent verification standards specifically for safety/alignment research quality.
+3 growthECAISA demands independent verification and worst-case epistemic norms that strengthen evaluation science for alignment.
+3 growthTaxonomy vocabulary is explicitly pitched for technical documentation, model-change tracking, and governance.
+3 growthAffective AI safety requires dedicated evaluation frameworks for cumulative, relational, and identity-level harms.
+3 growthThe AI Evaluability Gap concept reframes AI governance as requiring evidentiary foundations, not just system properties.
+3 growthYuvion VL's adversarially-aware pipeline advances the science of evaluating multimodal AI safety under realistic adversarial conditions.
+3 growthApothem-optimal robustness certifications provide more trustworthy safety guarantees for neural networks, advancing evaluation science.
+3 growthBenchmark-validity audit failures reveal the AI evaluability gap, showing governance assurance evidence can be silently manufactured.
+3 growthThe adversarial pragmatics benchmark provides linguistically controlled methodology for AI safety evaluation, advancing the science of AI evaluation.
+3 growthBehavioral divergence under quantization reveals that accuracy and perplexity metrics fail to capture important behavioral changes, informing evaluation science.
+3 growthModality and web-search conditions can reverse apparent safety/performance rankings used in deployment claims.
+3 growthUI/API and search confounds show safety claims need modality-aware evaluation science, not single-run API accuracy.
+3 growthUI/API modality and web-search conditions change measured safety scores, limiting naive benchmark-to-deployment claims.
+3 growthShows modality, web search, and repeated runs materially change safety-benchmark outcomes beyond single API accuracy.
+3 growthLanguage-slice gaps show collection-level multilingual safety benchmarks can overstate per-language protection.
+3 growthInteractive PCP offers a polynomial verifier path for honesty about probabilistic unwanted-outcome predictions.
+3 growthMaking understanding assessable extends evaluation science from systems to decision-maker epistemic adequacy.
+3 growthAnswer-access artifacts show process monitors can look better while missing critical wrong-trace correct-answer failures.
+3 growthPrivileged-access redesign shows prior internal-control claims may reflect superficial prompt-inferable targets rather than metacognition.
+3 growthCoverage-gap argument limits claims that cheaper automated red-teaming can retire human/contextual evaluators.
+3 growthDefines threat-model coverage gap showing benchmark red-teaming completeness is not deployment-context coverage.
+3 growthYuvion VL treats safety as an adversarial multimodal problem, advancing evaluation science for content safety.
+2 growthThe Inattentional Gap shows that benchmark safety scores decouple from real-world safety, challenging the science of AI evaluation.
+2 growthEvalSafetyGap framework contributes to the science of AI evaluation by formalizing the gap between evaluation proxies and alignment properties.
+2 growthFinding that evaluation-awareness shifts to earlier layers at scale complicates benchmark validity in AI evaluation science.
+2 growthEvalSafetyGap contributes to the science of AI evaluation by identifying systematic gaps between benchmark scores and latent safety properties.
+2 growthStrategic red teaming as a governance instrument advances the science of AI evaluation by making strategic uncertainty inspectable at the board level.
+2 growthSingle-prover interactive proofs provide a formal verification alternative to debate for AI safety evaluation.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
In June 2025 the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI), pivoting toward security, standards and adversary-model assessment.
NIST →Anthropic activated its ASL-3 deployment and security standard with Claude Opus 4 on 22 May 2025 — the first real-world trigger of a responsible-scaling tier, focused on blocking bio-weapon uplift.
Anthropic →The International Network of AI Safety Institutes (launched Nov 2024) ran a third joint testing exercise focused on agentic AI systems across cyber and fraud strands.
European Commission — AI Office →