← Back to the trend map

Safety · Concept

LLM Lie Detection and Deception Auditing

Study compares Diff-in-Means and INLP interventions for steering refusal in safety fine-tuned models, finding INLP counterfactual flipping competitive with directional ablation.

Trend strength 4/10
Momentum +4/q
Confidence medium
Status new
Forecast horizon

Verified model organisms and causal deception testbeds needed before lie-detection can be used in production auditing.

Connections

Connections · 1

How this node ties into the rest of the map, and the evidence behind each link.

Signal sources

Signal sources

Dated facts from primary sources in this direction.