← Back to the trend map

Safety · Trend

Strategic Attack Selection in Agentic AI Control Evaluations

Study of AI coding sabotage finds influential decision points broadly distributed throughout generated sequences, not concentrated at tool-call boundaries.

Trend strength 5/10
Momentum +5/q
Confidence medium
Status new
Forecast horizon

AI control evaluation frameworks must incorporate strategic attacker models to produce credible safety guarantees for deployment.

Connections

Connections · 2

How this node ties into the rest of the map, and the evidence behind each link.

Signal sources

Signal sources

Dated facts from primary sources in this direction.