Capabilities · Actor
Anthropic
Frontier lab shipping Claude Fable/Mythos 5.1, proposing lab-pace metrics, threat intel on misuse, and life-sciences verification.
Industry-wide jailbreak severity framework with Amazon, Microsoft, Google could become a de facto standard.
Connections
Connections · 25
How this node ties into the rest of the map, and the evidence behind each link.
Anthropic introduced Claude Sonnet 5 delivering frontier performance across coding, agents, and professional work.
+8 growthAnthropic proposed an industry-wide jailbreak severity scoring framework together with Amazon, Microsoft, Google, and Glasswing partners.
+8 growthAnthropic newsroom introduces Claude Opus 5 for long-running agents and professional work.
+8 growthAnthropic launched Claude Science as a customizable AI workbench for scientists.
+7 growthAnthropic case study covers UST bringing Claude to physical AI.
+7 growthAnthropic product lineup includes Claude for Teachers.
+7 growthAnthropic appointed Ben Bernanke to its Long-Term Benefit Trust governance body.
+6 growthAnthropic opened a research preview of the Model Hardware Standard for safe physical-device operation by AI agents.
+6 growthAnthropic developed the Jacobian lens tool providing the clearest glimpse yet at what is happening inside LLMs.
+5 growthAnthropic opens a research preview of the Model Hardware Standard for agents operating physical devices.
+5 growthSeptember 2026 Anthropic report details disrupted malicious Claude operations over eight months.
+5 growthAnthropic developed and deployed Fable 5 and Mythos 5 as frontier models.
+4 growthAnthropic opened a Seoul office and announced new partnerships across the Korean AI ecosystem.
+4 growthAnthropic newsroom introduces Claude Opus 5 for long-running agents and coding.
+4 growthBroader AI-for-science medicine design wave aligns with frontier labs’ science tooling and grant programs.
+4 growthAnthropic appoints Mariano-Florentino Cuéllar Chief Global Affairs Officer.
+4 growthAnthropic reports improving Fable 5 biology safeguards in an August product update.
+4 growthAnthropic explains how Claude’s text watermark works.
+4 growthCompetitive small open models trained on permissible data pressure proprietary labs’ openness and data-governance narratives.
+4 growthClaude-Opus-4.6 was one of four frontier models tested in the AI coding sabotage study.
+3 growthFrontier models including those from Anthropic were evaluated in the CogManip benchmark for manipulation risk.
+3 growthAnthropic newsroom highlights UST bringing Claude into physical AI deployments.
+3 growthThe paper evaluates Constitutional AI, a method developed by Anthropic, in the context of virtue ethics and existential risk.
+2 growthAnthropic is backing the respiratory infection research initiative alongside Stripe and OpenAI.
+2 growthAnthropic is one of the backers of the new respiratory infection research initiative alongside Stripe and OpenAI.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
The length of software tasks AI agents can do autonomously at 50% reliability has doubled about every 7 months — and since 2024 closer to every ~3 months.
METR →In one year scores rose by 18.8, 48.9 and 67.3 points on MMMU, GPQA and SWE-bench; real-world software solve rate jumped from 4.4% to 71.7%.
Stanford HAI — AI Index 2025 →On SWE-bench Verified (500 real GitHub issues), autonomous coding agents reached ~80–86% by late 2025, up from under 50% in early 2025.
Epoch AI →