← Back to the trend map

Safety · Concept

Item Response Theory for AI Safety Benchmarks

Psychometric IRT fit on eight safety benchmarks and 192 models yields refusal/truthfulness/harm factors, cheap adaptive testing, and sandbagging audits.

Trend strength 5/10
Momentum +5/q
Confidence medium
Status new
Forecast horizon

Adaptive IRT item sets could cut eval cost 97–99% and flag sandbagging.

Connections

Connections · 4

How this node ties into the rest of the map, and the evidence behind each link.

Signal sources

Signal sources

Dated facts from primary sources in this direction.