← Back to the trend map

Safety · Concept

Adversarial Pragmatics Benchmark for AI Safety Evaluation

Linguistically controlled benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, and policy ambiguity.

Trend strength 4/10
Momentum +4/q
Confidence medium
Status new
Forecast horizon

Connections

Connections · 3

How this node ties into the rest of the map, and the evidence behind each link.

Signal sources

Signal sources

Dated facts from primary sources in this direction.