Governance · Concept
AI Evaluability Gap
Benchmark-validity audits are themselves fragile and can be silently manufactured by implementation details, undermining governance assurance evidence.
Goodhart's Law dynamics in AI evaluation will intensify as optimization pressure on benchmarks grows.
Connections
Connections · 13
How this node ties into the rest of the map, and the evidence behind each link.
Sparse published audits in the Global South instantiate the broader evaluability gap under rapid deployment.
+4 growthThe AI Evaluability Gap concept reframes AI governance as requiring evidentiary foundations, not just system properties.
+3 growthBenchmark-validity audit failures reveal the AI evaluability gap, showing governance assurance evidence can be silently manufactured.
+3 growthProposed specification layer is positioned as connective tissue making evaluation, mediation, and escalation composable for oversight.
+3 growthTaxonomy aims to standardize how trained-model modifications are described for documentation and governance.
+3 growthShared vocabulary for how models were adapted improves documentation, change tracking, and evaluability for governors.
+3 growthCapability-based scenario rating is proposed because likelihood-based AI risk prediction is conceded to fail.
+3 growthMethodology makes decision-maker understanding explicit so safety cases alone do not fake evaluability.
+3 growthWhen evaluators disappear, post-hoc reconstruction assumptions fail, widening the evaluability/assurance gap.
+3 growthShort model-retirement intervals and version lock-in undermine assumptions that consequential AI decisions remain re-testable.
+3 growthWhen evaluators disappear, reconstructability of model-mediated decisions fails unless commitment and independent verifiability are designed in.
+3 growthAviation certification analysis reveals a structural gap in AI governance documents analogous to the AI evaluability gap.
+2 growthAffective safety harms are cumulative and relational, making them poorly captured by existing evaluation frameworks.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
EU AI Act obligations for general-purpose AI models applied from 2 Aug 2025; high-risk obligations under Annex III apply from 2 Aug 2026.
EU AI Act — implementation tracker →The Council of Europe Framework Convention on AI — the first legally binding international AI treaty — opened for signature in Sep 2024; the EU ratified it on 15 May 2026.
Council of Europe →On 26 Aug 2025 the UN General Assembly created an Independent International Scientific Panel on AI (40 experts) and a Global Dialogue on AI Governance.
United Nations →