Capabilities · Concept
DeepSeek-V4: Million-Token MoE Language Model
Open-weight model family whose cyber performance AISI finds closing the gap with frontier closed models.
Open-weight cyber parity timelines will keep compressing, stressing export and eval regimes.
Connections
Connections · 3
How this node ties into the rest of the map, and the evidence behind each link.
AISI cyber capability evals compare DeepSeek V4-Pro open weights to earlier frontier closed models.
+7 growthAISI cites DeepSeek V4-Pro among open models performing near frontier closed cyber capability.
+3 growthDeepSeek-V4's MoE architecture with million-token context enables new approaches to continual and long-context learning in LLMs.
+2 growthSignal sources
Signal sources
Dated facts from primary sources in this direction.
The length of software tasks AI agents can do autonomously at 50% reliability has doubled about every 7 months — and since 2024 closer to every ~3 months.
METR →In one year scores rose by 18.8, 48.9 and 67.3 points on MMMU, GPQA and SWE-bench; real-world software solve rate jumped from 4.4% to 71.7%.
Stanford HAI — AI Index 2025 →On SWE-bench Verified (500 real GitHub issues), autonomous coding agents reached ~80–86% by late 2025, up from under 50% in early 2025.
Epoch AI →