Anthropic August 2026 risk report raises misalignment to low, shelves Model 2
Anthropic published its Redacted Risk Report: August 2026 (Aug 14; coverage date July 15 under RSP v3.4), raising catastrophic-misalignment risk in high-stakes settings from “very low” to “low” amid uncertainty after cybersecurity-evaluation incident disclosures—while arguing the underlying case still likely supports “very low.” The report discloses unreleased internal Model 2 (somewhat more capable than Mythos 5; heavily used internally for coding/agents; no current external-release plan and incomplete predeployment assessments), notes Claude now authors a large majority of Anthropic’s merged production code with early R&D acceleration short of a 2× factor, and discloses that from May 2025–April 2026 all human-feedback vendor traffic (~50,000 contractors; ~133M exchanges) ran without blocking biological classifiers (since remediated; no CB misuse found). Distinct from the Summer 2026 agentic-misalignment case studies and from Claude text-watermark posts.




