Safety · Concept
Institutional Red-Teaming: Deployment Rules as Safety Variables
Treats institutional deployment rules and authority states as first-class safety variables for multi-agent systems.
Shifts safety R&D from model-only fixes toward tunable deployment institutions.
Connections
Connections · 5
How this node ties into the rest of the map, and the evidence behind each link.
IABench-CA shows deployment rules causally alter collective safety in multi-agent AI settings.
+4 growthPOLIS experiments show deployment rules, constitutional prompts, and executable guards reshape multi-agent violation rates.
+3 growthMulti-turn multi-agent stress tests treat deployment dialogue rules and strategies as first-class safety variables.
+3 growthBrings historical incident retrieval into pre-task planning, treating deployment safety process as an operational variable.
+3 growthInstitutional red-teaming methodology provides a concrete evaluation tool for the governance mechanisms identified in the agentic AI governance review.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
In June 2025 the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI), pivoting toward security, standards and adversary-model assessment.
NIST →Anthropic activated its ASL-3 deployment and security standard with Claude Opus 4 on 22 May 2025 — the first real-world trigger of a responsible-scaling tier, focused on blocking bio-weapon uplift.
Anthropic →The International Network of AI Safety Institutes (launched Nov 2024) ran a third joint testing exercise focused on agentic AI systems across cyber and fraud strands.
European Commission — AI Office →