Skip to content

Latest commit

 

History

History
575 lines (421 loc) · 34.3 KB

File metadata and controls

575 lines (421 loc) · 34.3 KB

Guide 03: Model Selection

The single highest-impact optimization. Choosing the right model per task can reduce your Claude Code bill by 30-60% with zero loss in output quality.

Most developers default to the most capable model for everything. This is like hiring a senior architect to change a lightbulb. Claude Fable 5 is extraordinary, but for renaming a variable it is a $50/M-output-token lightbulb-changer (and Opus 5 a $25 one). With the 4.7-generation tokenizer using up to 35% more tokens for the same text, the effective cost penalty for over-using top-tier models is even higher than posted pricing suggests.

Updated 2026-07-25 for the Opus 5 launch. Opus 5 (claude-opus-5, GA 2026-07-24) is the new Opus-tier flagship at the same $5/$25 as Opus 4.8, so the upgrade is free at the posted rate. Opus 4.8 is now legacy. Two behavior changes matter for cost: thinking is on by default (reasoning tokens bill as output at $25/1M), and the minimum cacheable prompt drops to 512 tokens from 1,024. See Guide 08 for the caching side and the Opus 5 notes below for the rest.


Table of Contents


Model Lineup and Pricing

Current Pricing (verified 2026-07-25, per 1M tokens)

Model Input Cost Output Cost Cache Hit 5m Cache Write 1h Cache Write Min cacheable prompt Relative Cost Context Window Max Output
Fable 5 (most capable) $10.00 $50.00 $1.00 $12.50 $20.00 512 2x baseline 1M 128K
Opus 5 (Opus flagship) $5.00 $25.00 $0.50 $6.25 $10.00 512 1x (baseline) 1M 128K
Opus 4.8 (legacy) $5.00 $25.00 $0.50 $6.25 $10.00 1,024 1x (baseline) 1M 128K
Opus 4.7 $5.00 $25.00 $0.50 $6.25 $10.00 2,048 1x (baseline) 1M 128K
Opus 4.6 $5.00 $25.00 $0.50 $6.25 $10.00 4,096 1x (baseline) 1M 128K
Opus 4.5 $5.00 $25.00 $0.50 $6.25 $10.00 4,096 1x (baseline) 200K 64K
Opus 4.1 $15.00 $75.00 $1.50 $18.75 $30.00 1,024 3x baseline 200K 32K
Sonnet 5 (Sonnet flagship) $3.00 $15.00 $0.30 $3.75 $6.00 1,024 ~1.67x cheaper 1M 128K
Sonnet 4.6 $3.00 $15.00 $0.30 $3.75 $6.00 1,024 ~1.67x cheaper 1M 64K
Sonnet 4.5 $3.00 $15.00 $0.30 $3.75 $6.00 1,024 ~1.67x cheaper 200K 64K
Haiku 4.5 $1.00 $5.00 $0.10 $1.25 $2.00 4,096 5x cheaper 200K 64K

Min cacheable prompt is the number of tokens a prefix must reach before cache_control does anything. Below the threshold the block is silently ignored: no error, no discount, full input price every turn. Opus 5 halving it to 512 means short system prompts that never cached on Opus 4.8 now do.

Opus 5 (GA 2026-07-24): the Opus-tier flagship and Anthropic's recommended starting point for complex agentic coding and enterprise work. Same $5/$25 as Opus 4.8, so the upgrade is free at the posted rate. 1M context at standard rates, 128K max output (300K on Batch via the output-300k-2026-03-24 beta), Batch $2.50/$12.50, Fast Mode supported at 2x ($10/$50), knowledge cutoff May 2026, earliest retirement 2027-07-24. Four cost-relevant changes versus Opus 4.8:

  1. Adaptive thinking is ON by default when you omit the thinking param. Reasoning tokens bill as output at $25/1M, and max_tokens caps thinking plus text -- carry over a small max_tokens and thinking can eat the budget before the answer is written. Raise it to 64K+ at xhigh/max effort.
  2. thinking: {type: "disabled"} is only legal at effort high or below. Pairing it with xhigh or max returns a 400. On Opus 4.8 that combination was allowed.
  3. Minimum cacheable prompt drops to 512 tokens (from 1,024).
  4. Cybersecurity safety classifiers ship with it. A cyber refusal can auto-fall-back to Opus 4.8 via the server-side fallbacks param (beta server-side-fallback-2026-07-01).

Opus 5 also writes longer output than Opus 4.8 by default and self-verifies its own work. Re-tune verbosity instructions after migrating, and delete carried-over "double-check your work" prompts -- you are now paying twice for behavior the model already performs.

Fable 5 (GA 2026-06-09): Anthropic's most capable widely released model -- a Mythos-class tier above Opus at $10/$50, 2x Opus 5. 1M context at standard rates, 128K max output, always-on adaptive thinking (control depth with effort), Batch supported ($5/$25), no Fast Mode, 30-day data retention required, min cacheable prompt 512. Safety classifiers can decline a request (HTTP 200 + stop_reason: "refusal"; pre-output refusals are free; beta fallbacks retries another model server-side). Cost guidance: reach for Fable 5 only when the task genuinely needs frontier-plus capability -- the hardest reasoning, the longest autonomous runs. For everything else Opus 5 at half the price is the efficient frontier. Mythos 5 is the same model minus the classifiers, limited to approved Project Glasswing customers.

Opus 4.8 status: Legacy as of the Opus 5 launch. Earliest retirement 2027-05-28. Same $5/$25 as Opus 5, so there is no cost argument for staying. Fast Mode supported at 2x. Min cacheable prompt 1,024. Reasons to pin it: prompts tuned to this snapshot, or a workload that needs thinking off at xhigh/max effort (which Opus 5 rejects with a 400). It also remains the server-side fallback target for Opus 5 cyber refusals.

Opus 4.7 status: Legacy. Earliest retirement 2027-04-16. Same price. Fast Mode was removed here -- speed: "fast" now returns an error with no fallback to standard. Min cacheable prompt 2,048. Pick 4.7 only if you have prompts pinned to that snapshot.

Opus 4.6 status: Legacy. Earliest retirement 2027-02-05. Same price. Fast Mode was removed here too, but silently -- speed: "fast" is accepted and runs at standard speed and standard rates (usage.speed comes back "standard"). Min cacheable prompt 4,096, the highest of any Opus. Pick 4.6 only if you want a stable snapshot or your prompts already perform well on it.

Opus 4.5 (200K-only): Legacy until at least 2026-11-24. Same price as the rest of the Opus tier but smaller context window, no Fast Mode, and a 4,096-token cache floor. Generally migrate to Opus 5 unless your code is pinned to this snapshot.

Sonnet 5 (claude-sonnet-5, GA 2026-06-30): the current Sonnet flagship -- best combination of speed and intelligence, adaptive thinking (effort defaults to high on the Claude API and Claude Code), 1M context at standard rates, 128K max output, no Fast Mode, min cacheable prompt 1,024. Uses the newer tokenizer (~30% more tokens for the same text). Introductory pricing $2/$10 per MTok through 2026-08-31, then standard $3/$15 (table above shows standard rates). New default for most production work; Sonnet 4.6 is now legacy.

Sonnet 4.5 (200K-only): Legacy until at least 2026-09-29. Same price as Sonnet 4.6 but smaller context window. Migrate to Sonnet 5 if you need 1M.

Opus 4.1: Still available at older pricing ($15/$75) -- 3x more expensive than current Opus tiers. Deprecated 2026-06-05, retires 2026-08-05. No reason to use it unless you have a specific compatibility need -- migrate to Opus 5.

Claude Mythos Preview (Project Glasswing): superseded by Mythos 5 -- retires 2026-06-30. It was the invitation-only research preview ($25/$125) for defensive cybersecurity through Glasswing partners. Its successor Mythos 5 drops to $10/$50 (same as Fable 5) and remains Glasswing-only; the prediction that Mythos-class capabilities would reach a widely released model came true as Fable 5, which is GA for everyone.

What These Numbers Mean in Practice

A typical Claude Code turn involves roughly 2,000-5,000 input tokens and 500-3,000 output tokens. Here is what a single turn costs across models:

Scenario Input Tokens Output Tokens Opus 5* Sonnet 5 Haiku 4.5
Quick fix (small) 2,000 500 $0.023 $0.014 $0.005
Component creation (medium) 5,000 2,000 $0.075 $0.045 $0.015
Architecture analysis (large) 10,000 5,000 $0.175 $0.105 $0.035
Multi-file refactor (XL) 20,000 10,000 $0.350 $0.210 $0.070

*Opus 5 costs shown at posted rate ($5/$25, identical to Opus 4.8). Multiply by ~1.2-1.35 to account for the 4.7-generation tokenizer's higher token counts on the same text. On Opus 5, add the default-on thinking tokens too: they are billed as output, so a turn that produced 500 visible output tokens may bill 1,500+ at high effort. Lower effort or disable thinking for routine turns.

Key insight: Output tokens cost 5x more than input tokens across all models. Tasks that generate a lot of code (scaffolding, boilerplate, test suites) are where model selection has the most impact.

Prompt Caching and Real-World Costs

Claude Code uses prompt caching, which reduces the cost of repeated input tokens by 90%. After the first turn of a session, cached content (system prompt, CLAUDE.md, stable conversation history) costs far less. This means:

  • First turn of a session is the most expensive
  • Subsequent turns benefit heavily from caching
  • Model selection still matters because output tokens are never cached, and output is where most cost accumulates in code-generation tasks
  • Model selection also sets the cache floor. Your prefix has to clear the model's minimum cacheable length before any of this applies -- 512 tokens on Opus 5 and Fable 5, 1,024 on Opus 4.8 and every Sonnet, 4,096 on Haiku 4.5 and Opus 4.6/4.5. Routing a short-prompt task down to Haiku saves 5x on the token rate but can lose the cache discount entirely.

The 80/20 Rule of Model Selection

80% of your Claude Code tasks can be handled by a cheaper model. The remaining 20% genuinely benefit from Opus.

Here is a breakdown of a typical developer's daily Claude Code usage:

Typical daily task distribution:
├── 40%  Simple tasks (formatting, renames, lookups, simple edits)     → Haiku
├── 40%  Medium tasks (components, bug fixes, tests, docs)             → Sonnet
└── 20%  Complex tasks (architecture, multi-file refactors, debugging) → Opus

Cost Impact of the 80/20 Rule

Assume 50 tasks per day with an average cost of $0.10/task on Opus 5:

Strategy Daily Cost Monthly Cost (22 days) Savings
All Opus 5 $5.00 $110.00 --
80/20 split (Haiku/Sonnet for 80%) $2.00 $44.00 60%
Optimized split (40/40/20) $1.70 $37.40 66%

Opus 5 posts the same $5/$25 as Opus 4.8 and 4.7/4.6, so the absolute dollar savings are smaller than they were against Opus 4.1 ($15/$75) -- but a 60-66% reduction still adds up fast across a team. Model selection remains the single biggest lever you have, and two things make skipping Opus more valuable than the table suggests: the tokenizer overhead, and Opus 5's default-on thinking, which inflates the output side of every turn you route there unnecessarily.


Task Complexity Decision Tree

Use this decision tree to quickly decide which model to use:

START: What is the task?
│
├── Does it require understanding complex architecture or
│   reasoning about multi-file interactions?
│   ├── YES → Is it a planning/analysis task (no code output)?
│   │         ├── YES → Opus 5 (plan mode)
│   │         └── NO  → Opus 5
│   └── NO ↓
│
├── Does it require generating new code with non-trivial logic?
│   ├── YES → Is the logic self-contained in 1-2 files?
│   │         ├── YES → Sonnet
│   │         └── NO  → Sonnet (consider Opus if > 5 files)
│   └── NO ↓
│
├── Is it a mechanical/repetitive change?
│   (formatting, renaming, simple find-replace, adding imports)
│   ├── YES → Haiku
│   └── NO ↓
│
├── Is it a lookup or question about the codebase?
│   ├── YES → Haiku (for simple questions) / Sonnet (for analysis)
│   └── NO ↓
│
└── Default → Sonnet (the safest general-purpose choice)

The "Two Question" Shortcut

If the decision tree feels heavy, just ask two questions:

  1. Does this task require reasoning about how multiple components interact? If yes: Opus. If no: continue.
  2. Does this task require generating non-trivial new logic? If yes: Sonnet. If no: Haiku.

Task Categories with Recommended Models

Simple Tasks: Use Haiku 4.5

Cost per task: $0.003-$0.02

Haiku handles these tasks with the same quality as more expensive models. There is no benefit to using Sonnet or Opus here.

Task Example Why Haiku Works
Code formatting "Fix the indentation in utils.py" Mechanical transformation, no reasoning needed
Variable/function renaming "Rename getData to fetchUserProfile" Simple find-and-replace with scope awareness
Import management "Add missing imports to this file" Pattern matching against existing code
Simple type annotations "Add TypeScript types to these function params" Inferring types from usage patterns
Comment updates "Update the JSDoc for this function" Reading function signature, writing description
Config file changes "Add cors: true to the server config" Small edits to structured files
Git operations "Create a commit message for these changes" Summarizing diffs
File lookups "What files import from utils/auth?" Grep-based search, no deep reasoning
Simple error fixes "Fix this missing semicolon / closing bracket" Syntax-level corrections
Moving/copying files "Move Header.tsx to components/layout/" File system operations with import updates

How to invoke:

claude --model haiku "rename getUserData to fetchUserProfile in src/api/"

Medium Tasks: Use Sonnet 5

Cost per task: $0.02-$0.15

Sonnet is the sweet spot for most development work. It handles logic, generates quality code, and understands context well.

Task Example Why Sonnet Works
Component creation "Create a pagination component with prev/next" Generates well-structured code with standard patterns
Bug fixes "Fix the race condition in the auth flow" Understands cause-and-effect in code, traces logic
Test writing "Write unit tests for the CartService class" Follows testing patterns, covers edge cases
API endpoint creation "Add a PUT endpoint for updating user profiles" Follows existing patterns in the codebase
Database queries "Write a query to get users with expired subs" Understands schema relationships
Documentation "Write API docs for the payment module" Reads code, generates structured documentation
Code review "Review this PR diff for issues" Identifies common anti-patterns and bugs
Refactoring (single file) "Extract the validation logic into its own function" Restructures code while preserving behavior
Error handling "Add proper error handling to the API layer" Understands failure modes, generates try/catch patterns
State management "Add Redux slice for the notification feature" Follows established state patterns in the project

How to invoke:

claude --model sonnet "write unit tests for src/services/CartService.ts"

Complex Tasks: Use Opus 5

Cost per task: $0.05-$0.50+ (higher in practice due to the tokenizer and default-on thinking)

Reserve Opus for tasks where deep reasoning, multi-file coordination, or architectural understanding provides genuine value. At $5/$25 (same as 4.8/4.7/4.6, down from 4.1's $15/$75) the cost penalty for using it is smaller than it used to be, but it is still ~1.67x more than Sonnet and 5x more than Haiku at posted rates, the tokenizer adds another 20-35% of effective cost, and thinking-on-by-default adds billable output on top of that. Defaulting to Opus for every task remains wasteful.

Why Opus 5 over Opus 4.8: same posted price, better agentic coding, and a cache floor half as high (512 versus 1,024 tokens), so more of your system prompt actually caches. Opus 5 also self-verifies its work before reporting and takes instructions literally, which on long autonomous runs (audit, multi-file refactor, migration) often pays for itself by avoiding retry loops. The flip side is verbosity: Opus 5 writes longer than 4.8 by default, so trim your verbosity instructions rather than assuming an old prompt is still cost-optimal.

Thinking modes differ, and this is the migration trap:

  • Opus 5 uses adaptive thinking, ON by default when the thinking param is omitted. Effort defaults to high. thinking: {type: "disabled"} works only at effort high or below -- pairing it with xhigh or max is a 400.
  • Opus 4.8 uses adaptive thinking, off unless requested. Effort defaults to high on all surfaces (the Claude Code default was xhigh on 4.7).
  • Opus 4.7 uses adaptive thinking plus an xhigh effort level.
  • Opus 4.6 uses extended thinking -- you can configure a reasoning token budget.
  • If you have prompts or harnesses tuned against extended thinking's explicit budget knobs, they will not carry over to 4.8 or 5. And if you are moving from 4.8 to 5, the same request now costs more unless you explicitly manage thinking: budget the difference before rolling out, and raise max_tokens to 64K+ if you run xhigh or max.

Benchmarks published by Anthropic for Opus 4.6 (the numbers were even higher for Mythos Preview, the tier that Fable 5 / Mythos 5 now succeed — included below for reference):

Benchmark Opus 4.6 Mythos Preview (retired tier)
SWE-bench Verified 80.8% 93.9%
SWE-bench Pro 53.4% 77.8%
Terminal-Bench 2.0 65.4% 82.0%
CyberGym (vuln reproduction) 66.6% 83.1%

Mythos-class capability is now generally available as Fable 5 ($10/$50). Anthropic has not published a per-benchmark scorecard for Fable 5 or Opus 5 on these exact suites at the time of writing (2026-07-25) -- when official numbers land, this table should be extended.

Task Example Why Opus Is Worth It
Architecture design "Design the module structure for a plugin system" Requires reasoning about abstractions, trade-offs, extensibility
Multi-file refactoring "Migrate from REST to GraphQL across 15 files" Needs to hold the full picture, coordinate changes, avoid breakage
Complex debugging "Find why checkout fails intermittently under load" Requires reasoning across multiple systems, race conditions, state
Performance optimization "Identify and fix the N+1 queries in the API" Needs to trace data flow through multiple layers
Security auditing "Review the auth system for vulnerabilities" Requires deep understanding of attack vectors, subtle bugs
Migration planning "Plan the migration from Webpack to Vite" Needs to understand build system internals, dependency implications
Design pattern implementation "Implement CQRS for the order processing system" Complex pattern with many interacting components
System integration "Integrate Stripe webhooks with our event system" Multiple systems, error handling, idempotency concerns
Algorithm development "Implement a rate limiter with sliding window" Algorithmic reasoning, edge cases, correctness proofs
Legacy code understanding "Explain how the billing engine works end-to-end" Reading and synthesizing across a large, undocumented codebase

How to invoke:

claude --model opus "design a plugin architecture for our CLI tool"

Want a specific Opus snapshot explicitly? Use the dated/aliased ID:

  • Opus 5: --model claude-opus-5 (Claude API), anthropic.claude-opus-5 (Bedrock Messages API), us.anthropic.claude-opus-5 (Bedrock legacy InvokeModel/Converse), claude-opus-5 (Vertex AI)
  • Opus 4.8: --model claude-opus-4-8 (Claude API), anthropic.claude-opus-4-8 (Bedrock Messages API), us.anthropic.claude-opus-4-8 (Bedrock legacy InvokeModel/Converse)
  • Opus 4.7: --model claude-opus-4-7 (Claude API), anthropic.claude-opus-4-7 (Bedrock)
  • Opus 4.6: --model claude-opus-4-6 (Claude API), anthropic.claude-opus-4-6-v1 (Bedrock legacy InvokeModel/Converse)
  • Opus 4.5: --model claude-opus-4-5-20251101
  • Opus 4.1: --model claude-opus-4-1-20250805 (retires 2026-08-05)

The opus alias maps to Opus 5 on current Claude Code releases. If you need a pinned snapshot, name it explicitly rather than relying on the alias -- the alias moves with each Opus launch.


How to Set the Model

Method 1: The --model Flag (Per-Task)

The most direct approach. Override the model for a single invocation:

# Use Haiku for a quick rename
claude --model haiku "rename processData to transformPayload in src/"

# Use Sonnet for a component
claude --model sonnet "create a modal dialog component"

# Use Opus for architecture work
claude --model opus "design the caching layer for our API"

Best for: Ad-hoc tasks where you know the complexity upfront.

Method 2: Default Model in Settings (Session Default)

Set a default model in your Claude Code settings (~/.claude/settings.json or project-level .claude/settings.json):

{
  "$schema": "https://json.schemastore.org/claude-code-settings.json",
  "model": "sonnet"
}

Then override with --model only when needed. This way, your baseline cost is Sonnet-level, and you opt into Opus explicitly.

Best for: Establishing a cost-efficient baseline across all sessions.

Method 3: Command Frontmatter (Per-Command)

Define the model inside a Claude Code custom command (.claude/commands/*.md):

---
model: haiku
---

Fix formatting issues in the files I specify. Do not change logic or behavior.
Only fix indentation, trailing whitespace, and missing newlines.

This ensures the command always uses the specified model regardless of your session default.

Best for: Repetitive tasks that should always use a specific model.

Method 4: Subagent Model Configuration

When using the Task tool to delegate work to a subagent, the subagent uses its own model configuration. You can instruct the main agent to delegate specific tasks to cheaper subagents:

Use a subagent to find all files that import from 'utils/deprecated'.
The subagent should use Haiku since this is a simple search task.

You can also configure subagent model preferences in your CLAUDE.md:

# Subagent Guidelines
- Use Haiku for: file searches, simple edits, formatting
- Use Sonnet for: code generation, test writing, bug fixes
- Use Opus only for: architecture decisions, complex multi-file work

Best for: Automated delegation where different subtasks have different complexity levels.


Cost-Per-Task Examples

These are real-world estimates based on typical token usage patterns. All costs assume prompt caching is active (not the first turn of a session).

The Opus 5 and Sonnet 5 columns use the same posted rates their predecessors did ($5/$25 and $3/$15), so these figures carry over unchanged from the Opus 4.8 / Sonnet 4.6 era. What they do not include is Opus 5's default-on thinking: reasoning tokens bill as output at $25/1M, so an Opus 5 turn at high effort can bill noticeably more than the output column shows. Treat the Opus 5 numbers as a floor for thinking-off or low-effort runs.

Example 1: Rename a Function

Task: Rename getUserData to fetchUserProfile across 8 files.

Haiku 4.5 Sonnet 5 Opus 5
Input tokens ~3,000 ~3,000 ~3,000
Output tokens ~1,200 ~1,200 ~1,200
Cost $0.009 $0.027 $0.045
Quality Identical Identical Identical

Verdict: Haiku. Saves $0.036 per rename vs Opus. Over 10 renames/day, that is $0.36/day saved. The gap is narrower than it used to be with old Opus pricing, but Haiku is still 5x cheaper for tasks where quality is identical.

Example 2: Write Unit Tests for a Service

Task: Write comprehensive unit tests for PaymentService (5 methods, ~200 lines).

Haiku 4.5 Sonnet 5 Opus 5
Input tokens ~8,000 ~8,000 ~8,000
Output tokens ~4,000 ~4,000 ~4,000
Cost $0.028 $0.084 $0.140
Quality Good (may miss edge cases) Very good Excellent (not worth 1.67x more)

Verdict: Sonnet. Saves $0.056 vs Opus with negligible quality difference for standard test writing.

Example 3: Debug a Race Condition

Task: Investigate and fix intermittent auth failures in a distributed system spanning 12 files.

Haiku 4.5 Sonnet 5 Opus 5
Turns needed ~15 (struggles) ~8 ~4
Total input tokens ~120,000 ~80,000 ~60,000
Total output tokens ~30,000 ~25,000 ~20,000
Total cost $0.270 $0.615 $0.800
Quality Poor (likely fails) Decent High (finds root cause)
Time 25 min 15 min 8 min

Verdict: Opus. While the cost gap between Opus and Sonnet is now much narrower ($0.80 vs $0.615), Opus solves it in fewer turns and developer time saved justifies the modest premium. This is also the task type where Opus 5's default-on thinking earns its keep: race-condition debugging is exactly the reasoning-heavy work those tokens buy, and the fewer-turns effect more than offsets them.

Example 4: Create a CRUD API Endpoint

Task: Add a new /api/projects endpoint with GET, POST, PUT, DELETE.

Haiku 4.5 Sonnet 5 Opus 5
Input tokens ~6,000 ~6,000 ~6,000
Output tokens ~3,500 ~3,500 ~3,500
Cost $0.024 $0.071 $0.118
Quality Adequate (follows patterns) Good Excellent (overkill)

Verdict: Sonnet. Standard CRUD follows patterns that Sonnet handles well.

Example 5: Plan a Database Migration

Task: Design the migration strategy from MongoDB to PostgreSQL for a 20-collection database with complex relationships.

Haiku 4.5 Sonnet 5 Opus 5
Input tokens ~15,000 ~15,000 ~15,000
Output tokens ~8,000 ~8,000 ~8,000
Cost $0.055 $0.165 $0.275
Quality Superficial plan Good plan Thorough, catches edge cases

Verdict: Opus. Migration planning is high-stakes. A missed edge case can cost days of developer time. The extra $0.11 vs Sonnet is negligible compared to the cost of a botched migration.


The Common Mistake: Opus for Everything

The "Premium Default" Anti-Pattern

Many developers set Opus as their default model and never change it. Their reasoning: "I want the best output, and the cost is acceptable."

With Opus 5's pricing ($5/$25, unchanged from 4.8/4.7/4.6, down from 4.1's $15/$75), this anti-pattern is less financially devastating than it used to be, but it is still wasteful, and two things quietly widen the gap back out: the 4.7-generation tokenizer (up to 35% more tokens per turn) and Opus 5's default-on thinking (extra billed output on every turn, including the trivial ones).

1. Quality is often identical across models for simple tasks.

For formatting, renaming, import management, config changes, and other mechanical tasks, Haiku produces output that is indistinguishable from Opus. You are paying 5x more for the same result.

2. The cost still compounds.

Scenario: 50 tasks/day, 22 working days/month

All Opus 5:  50 tasks x $0.10 avg x 22 days = $110/month (+35% tokenizer = ~$148,
             more once default-on thinking is counted)
Smart split: (20 x $0.005) + (20 x $0.03) + (10 x $0.10) x 22 = $37/month

Annual difference: $876
Team of 5 annual difference: $4,380

While the savings are more modest than they were at old Opus pricing, $4,380/year for a 5-person team is still worth capturing — especially since it requires no loss in output quality.

3. More expensive does not mean faster for simple tasks.

Opus does not rename a variable faster than Haiku. The latency is often higher because Opus generates more thorough (but unnecessary) reasoning for simple tasks -- and on Opus 5 that reasoning now happens by default, so the penalty applies unless you explicitly lower effort or disable thinking.

How to Break the Habit

  1. Set Sonnet as your default model. This is the best general-purpose starting point.
  2. Create Haiku commands for your most frequent simple tasks (formatting, renaming, lookups).
  3. Explicitly opt into Opus with --model opus only when you are doing architecture, complex debugging, or multi-file planning.
  4. Review your usage weekly. Look at which tasks used Opus and ask: "Did that task genuinely need Opus?"

Advanced: Dynamic Model Routing

CLAUDE.md-Based Routing Guidelines

Add model routing guidance to your CLAUDE.md so Claude Code itself helps you pick the right model when delegating to subagents:

# Model Routing
When delegating subtasks:
- Haiku: file searches, grep operations, simple edits, formatting, git operations
- Sonnet: code generation, test writing, bug fixes, documentation, single-file refactors
- Opus: architecture decisions, multi-file refactors, complex debugging, security reviews

Cost-Aware Command Library

Build a library of commands with pre-assigned models:

.claude/commands/
├── format.md          (model: haiku)
├── rename.md          (model: haiku)
├── find-usages.md     (model: haiku)
├── write-test.md      (model: sonnet)
├── fix-bug.md         (model: sonnet)
├── create-component.md (model: sonnet)
├── review-arch.md     (model: opus)
├── plan-refactor.md   (model: opus)
└── security-audit.md  (model: opus)

This way, model selection is built into your workflow. You do not have to think about it each time.

The Escalation Pattern

Start with the cheapest model and escalate only if needed:

  1. Try Haiku first. If the output is good, you are done.
  2. If Haiku struggles, retry with Sonnet.
  3. If Sonnet struggles, retry with Opus.

This sounds like it wastes tokens on failed attempts, but in practice, most tasks succeed on the first try with the cheaper model, and the few that need escalation still cost less overall than defaulting to Opus.


Quick Reference Card

HAIKU 4.5 ($1/$5 per 1M tokens)
├── Formatting and linting fixes
├── Variable and function renaming
├── Import management
├── Config file edits
├── File searches and lookups
├── Git commit messages
├── Simple type annotations
└── Mechanical find-and-replace

SONNET 5 ($3/$15 per 1M tokens, $2/$10 intro through 2026-08-31)
├── Component and module creation
├── Bug fixes (single file or simple multi-file)
├── Unit and integration test writing
├── API endpoint creation
├── Documentation generation
├── Code review
├── Single-file refactoring
└── Error handling implementation

OPUS 5 ($5/$25 per 1M tokens, +~35% tokenizer overhead, thinking ON by default)
├── Architecture design and planning
├── Multi-file refactoring (5+ files)
├── Complex debugging (race conditions, memory leaks)
├── Long autonomous agentic runs (self-verifies outputs)
├── Performance optimization (N+1 queries, bottlenecks)
├── Security auditing
├── Migration planning
├── System integration design
├── Legacy code comprehension
└── Cost control: lower `effort` for routine turns, or thinking {type:"disabled"}
    at effort high or below (xhigh/max + disabled = 400)

FABLE 5 ($10/$50 per 1M tokens -- 2x Opus 5, most capable widely released model)
├── The absolute hardest reasoning problems Opus 5 can't crack
├── The longest autonomous agentic runs (single turns can run many minutes)
├── Mythos-class capability without a Glasswing invitation
└── Note: always-on thinking, no Fast Mode, safety classifiers may refuse
    (pre-output refusals are free; use the beta fallbacks param to retry)

OPUS 4.8 ($5/$25 per 1M tokens -- legacy, same price as Opus 5, Fast Mode 2x)
└── Only if your prompts are tuned to this snapshot, or you need thinking off
    at xhigh/max effort (Opus 5 rejects that combination)

OPUS 4.7 ($5/$25 per 1M tokens -- legacy, Fast Mode REMOVED: speed "fast" errors)
└── Only if you have prompts pinned to that snapshot

OPUS 4.6 ($5/$25 per 1M tokens -- legacy, Fast Mode REMOVED: silently runs standard)
└── Only if you have prompts tuned to the older tokenizer or want a stable snapshot

FAST MODE -- OPUS 5 or OPUS 4.8 only ($10/$50 per 1M tokens -- 2x premium, beta,
Claude API + Managed Agents only; the old 6x tier no longer exists)
└── Only when output latency directly impacts revenue / UX
    (live demos, real-time agentic loops, urgent debugging)
    NOT for routine interactive coding

Next Steps

  • Set Sonnet as your default model today
  • Create 2-3 Haiku commands for your most common simple tasks
  • Read Guide 04: Workflow Patterns for additional cost-saving strategies
  • Use the Token Estimator to measure costs before and after switching models

Back to README | Previous: Context Optimization | Next: Workflow Patterns