Cited source trail
Claude Sonnet 5
Original question
Investigate this live topic: Claude Sonnet 5. Start with https://www.anthropic.com/news/claude-sonnet-5 and browse beyond it. Summarize what changed, why it matters, and cite the strongest sources.
# π Claude Sonnet 5 β Full Investigation Report
**Released:** June 30, 2026 | **Announcement:** [anthropic.com/news/claude-sonnet-5](https://www.anthropic.com/news/claude-sonnet-5)
---
## π§ What Changed (vs Sonnet 4.6)
Claude Sonnet 5 is a **massive leap forward** in agentic capability β Anthropic describes it as *"the most agentic Sonnet model yet."* It makes plans, uses browsers and terminals, self-corrects, and runs autonomously at a level that previously required the much pricier Opus-class models.
### Key Benchmark Improvements
| Benchmark | What It Measures | Sonnet 5 | Sonnet 4.6 | Ξ |
|---|---|---|---|---|
| **SWE-bench Pro** | Real-world software engineering | **63.2%** | 58.1% | +5.1 pp |
| **Terminal-Bench 2.1** | Command-line agent tasks | **80.4%** | 67.0% | **+13.4 pp** β first model to break 80% |
| **OSWorld-Verified** | Computer use (GUI control) | **81.2%** | 78.5% | +2.7 pp |
| **Humanity's Last Exam** | Graduate-level reasoning | **57.4%** w/tools | ~47% | **+10 pp** |
| **GDPval-AA v2** | Knowledge work quality | **1,618** | β | *Beats Opus 4.8 (1,615)* |
| **AA-Briefcase** | Multi-step knowledge work | *Matches/outperforms Opus 4.8* | β | *Trails only Fable 5* |
| **SciCode** | Scientific coding | β | β | **+7 pp** |
| **CritPt** | Frontier physics reasoning | 17% | β | +14 pp over 4.6 (but behind Opus) |
| **BrowseComp** | Agentic web search | *Matches Opus 4.8 at max effort* | β | Huge improvement |
### Architectural & Feature Changes
- **1M token context window** (same as Sonnet 4.6)
- **New "xhigh" effort setting** β Sonnet 5 now has 5 effort levels (low, medium, high, xhigh, max), matching Opus 4.8
- **Training cutoff:** January 2026
- **More token-hungry:** Uses ~40% more output tokens per Intelligence Index task than Sonnet 4.6, reflecting deeper reasoning
- **~3x agentic turns** in knowledge work evaluations (GDPval-AA, AA-Briefcase)
### Safety Improvements
- **Lower hallucination rate** than Sonnet 4.6
- **Lower sycophancy** than Sonnet 4.6
- **Better refusal** of malicious requests
- **Better resistance** to prompt injection hijack attempts
- **Lower cybersecurity capability** by design (deliberately not trained on cyber tasks β unlike current Opus models)
- Automated behavioral audit: overall safer than Sonnet 4.6, though not as clean as Opus 4.8/Mythos Preview
### Pricing
| Period | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| **Through Aug 31, 2026** (intro) | **$2** | **$10** |
| **After Aug 31, 2026** | $3 | $15 |
**Cache:** $3.75/M write (25% premium, 5-min TTL), $0.30/M cache hit (90% discount)
---
## β Why It Matters
### 1. Agentic capability hit a new price-performance frontier
Sonnet 5 brings **Opus-competitive agentic behavior** (autonomous planning, multi-step tool use, self-verification) to a mid-tier price point. Early testers report it *"carries multi-step pull requests through to a tested, verified result on its own"* β things previous Sonnet models would stall on or never finish.
### 2. Knowledge work parity with Opus 4.8
On real-world professional benchmarks like **GDPval-AA v2** and **AA-Briefcase**, Sonnet 5 matches or **slightly beats Opus 4.8** β the model that costs 67% more per token. This is a genuine breakthrough for the Sonnet line.
### 3. A nuanced cost story
**Artificial Analysis** found a catch: Sonnet 5's deeper agentic behavior means it uses ~40% more tokens per task, raising per-task cost to **$2.29** on the Intelligence Index β 2Γ Sonnet 4.6 and 15% *more* than Opus 4.8. So while per-token pricing is lower, the agentic "thinking tax" can offset savings on complex tasks. However, the introductory pricing ($2/$10) significantly mitigates this, and for many users the raw capability uplift justifies the cost.
### 4. Broader context: model regulation
Sonnet 5 launches alongside news that the U.S. government **blocked Claude Mythos 5 and Claude Fable 5** over cybersecurity concerns β models that exceeded critical capability thresholds in the U.S. AI Safety Institute's assessments. Sonnet 5, with deliberately reduced cyber capabilities, is the "safe" frontier model that can ship without export restrictions.
### 5. Pushes the entire ecosystem forward
Anthropic also published the Sonnet 5 System Card (PDF) in collaboration with **Mozilla** β a new standard for pre-deployment safety transparency. The model is already live across all plans (Free, Pro, Max, Team, Enterprise) and in Claude Code.
---
## π Intelligence Ranking (Artificial Analysis Index)
| Rank | Model | Score |
|---|---|---|
| #1 | Mythos Preview | β |
| #2 | Fable 5 | β |
| #3 | GPT-5.5 (Pro) | β |
| #4 | Opus 4.8 (max) | ~55β56 |
| **#5** | **Claude Sonnet 5 (max)** | **53** |
| ~#6 | GPT-5.5 (high) | ~55β56 |
Sonnet 5 is **only 2β3 points behind GPT-5.5 (xhigh)** and Opus 4.8 on the Intelligence Index.
---
## π Strongest Sources (Cited)
| Source | URL | Why It's Authoritative |
|---|---|---|
| **Anthropic (official)** | [anthropic.com/news/claude-sonnet-5](https://www.anthropic.com/news/claude-sonnet-5) | Primary source β official benchmarks, product details, safety evaluation summary |
| **Anthropic System Card** (PDF) | [anthropic.com/claude-sonnet-5-system-card](https://www.anthropic.com/claude-sonnet-5-system-card) | Full pre-deployment safety evaluation with Mozilla collab |
| **Artificial Analysis** | [article](https://artificialanalysis.ai/articles/claude-sonnet-5-agentic-cost) | Independent rigorous benchmarking with per-task cost analysis β worked with Anthropic pre-release on evaluation |
| **The Decoder** | [the-decoder.com](https://the-decoder.com/anthropics-new-claude-sonnet-5-closes-the-gap-to-the-pricier-opus-model-series/) | Independent journalism with detailed benchmark table |
| **TechCrunch** | [techcrunch.com](https://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/) | Major tech publication covering launch and market positioning |
1navigatehttps://www.anthropic.com/news/claude-sonnet-5
2clickundefined
3searchClaude Sonnet 5 review analysis benchmarks
4navigatehttps://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/
5navigatehttps://thenewstack.io/claude-sonnet-5-launch/
6navigatehttps://the-decoder.com/anthropics-new-claude-sonnet-5-closes-the-gap-to-the-pricier-opus-model-series/
7navigatehttps://artificialanalysis.ai/models/claude-sonnet-5
8navigatehttps://artificialanalysis.ai/articles/claude-sonnet-5-agentic-cost
9ddm--text
10navigatehttps://www.anthropic.com/news/claude-sonnet-5
11ddm--text
12searchClaude Sonnet 5 SWE-bench Terminal-Bench OSWorld benchmark scores