Safety · Concept
Model Hypnosis: Additive Subliminal Prompt Control
Phenomenon where weak inconspicuous prompt cues combine to strongly control model behavior across families and scales.
Creates new attack surface and interpretability obstacle for frontier systems.
Connections
Connections · 5
How this node ties into the rest of the map, and the evidence behind each link.
Additive inconspicuous prompt control creates stealthy jailbreak-like failures that severity frameworks may need to score.
+6 growthAdditive inconspicuous cues offer a distinct control channel related to jailbreak-style behavioral steering.
+5 growthHypnosis shows stealthy additive prompt cues can strongly control models, expanding beyond classic jailbreak framings.
+3 growthAugust arXiv safety batch includes the model hypnosis demonstration across frontier reasoning models.
+3 growthModel hypnosis paper appears in the day's arXiv AI safety cluster.
+3 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
In June 2025 the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI), pivoting toward security, standards and adversary-model assessment.
NIST →Anthropic activated its ASL-3 deployment and security standard with Claude Opus 4 on 22 May 2025 — the first real-world trigger of a responsible-scaling tier, focused on blocking bio-weapon uplift.
Anthropic →The International Network of AI Safety Institutes (launched Nov 2024) ran a third joint testing exercise focused on agentic AI systems across cyber and fraud strands.
European Commission — AI Office →