/model-routing-patterns
Multi-model pipelines (Haiku/Sonnet/Opus): cost routing, escalation, fallback chains. Triggers: model routing, Haiku, Sonnet, Opus, escalation, fallback chain.
$ npx -y skills add softspark/ai-toolkit --skill model-routing-patterns --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/model-routing-patterns
Context preview
The summary Claude sees to decide when to auto-load this skill.
Multi-model pipelines (Haiku/Sonnet/Opus): cost routing, escalation, fallback chains. Triggers: model routing, Haiku, Sonnet, Opus, escalation, fallback chain.
SKILL.md
model-routing-patterns.SKILL.mdname: model-routing-patterns
description: "Multi-model pipelines (Haiku/Sonnet/Opus): cost routing, escalation, fallback chains. Triggers: model routing, Haiku, Sonnet, Opus, escalation, fallback chain."
effort: medium
user-invocable: false
allowed-tools: Read
Model Routing Patterns
Three Claude tiers. Using Opus for everything is 10-40x more expensive than it needs to be. Using Haiku for everything loses accuracy on hard tasks. The craft is routing.
Model Characteristics (2026)
| Model | $/1M in·out | Cost (rel.) | Strengths | When | |-------|-------------|-------------|-----------|------| | Haiku 4.5 | $1 / $5 | 1x | Classification, extraction, simple tools, moderation | Bulk processing, triage, labels | | Sonnet 5 | $3 / $15 | ~3x | General coding, reasoning, most agent tasks | Default workhorse | | Opus 4.8 | $5 / $25 | ~5x | Complex reasoning, orchestration, architecture, large context | Hard, rare, high-stakes | | Fable 5 | $10 / $50 | ~10x | Most demanding long-horizon agentic work | Only when explicitly chosen |
Prices are per 1M tokens; ratios are approximate and shift between releases. Re-check pricing before committing a production path.
> **Fable 5 is not the default "best model".** Its price sits above Opus-tier, and Opus 4.8 is state-of-the-art on planning/orchestration at half the input and output cost. Reach for Fable 5 only when the user explicitly asks for it or a benchmarked task genuinely needs it — for "use the strongest model", the target is `claude-opus-4-8`.
Effort — the cheaper lever before swapping models
On Fable 5 / Opus 4.8 / Sonnet 5, `output_config.effort` (`low` | `medium` | `high` | `xhigh` | `max`) controls thinking depth and token spend **without changing the model** — so it does not invalidate the prompt cache the way a mid-session model swap does. Tune effort first; drop to a cheaper model only when effort alone can't hit the cost target.
| Effort | Use for | |--------|---------| | `low` | Latency-sensitive, non-intelligence-sensitive: chat, simple lookups, cheap subagents | | `medium` | Cost-conscious step-down from the default | | `high` | Default for most intelligence-sensitive work (a good quality/cost balance) | | `xhigh` | Hardest coding and agentic tasks (Claude Code's default) | | `max` | Correctness matters more than cost; test for diminishing returns |
In our agents, effort is set per skill/agent frontmatter (`effort:`), not swapped at runtime. Combine effort routing with model routing: e.g. `sonnet` at `high` often beats `opus` at `low` for cost-equal quality — benchmark before committing.
Pattern 1 — Complexity Router (pre-classify)
Cheap model classifies the request, then routes to the right tier:
def route(user_message: str) -> str:
complexity = classify_with_haiku(user_message) # returns: simple | medium | hard
return {"simple": "haiku", "medium": "sonnet", "hard": "opus"}[complexity]Good when ~60% of traffic is simple. Overhead: one Haiku call per request (~100 tokens).
Pattern 2 — Confidence-Based Escalation
Try the cheap model first, escalate only when it hesitates:
def solve(problem: str):
haiku = call_haiku(problem)
if haiku.confidence > 0.85:
return haiku.answer
sonnet = call_sonnet(problem + haiku.reasoning)
if sonnet.confidence > 0.8:
return sonnet.answer
return call_opus(problem)Haiku must be prompted to output confidence (e.g. via tool-use structured output — see `json-mode-patterns`). Pure self-reported confidence is noisy; combine with a heuristic (output length, tool calls, hedging words).
Pattern 3 — Sub-agent Delegation (Opus orchestrates, Haiku workers)
Orchestrator reasons about the plan, workers execute atomic steps:
Opus (planner)
├── Haiku (extract_dates_from_doc_1)
├── Haiku (extract_dates_from_doc_2)
├── Haiku (extract_dates_from_doc_3)
└── Opus (synthesize all extractions into timeline)
Real example: `/orchestrate` in ai-toolkit runs Opus as planner, subagents (model per agent's frontmatter) as workers. See `app/agents/*.md` — each agent sets `model:` explicitly.
Pattern 4 — Fallback Chain (resilience, not cost)
When primary is rate-limited or errors, degrade gracefully:
def call_with_fallback(messages):
for model in ["claude-opus-4-8", "claude-sonnet-5", "claude-haiku-4-5"]:
try:
return client.messages.create(model=model, messages=messages, ...)
except (RateLimitError, OverloadedError):
continue
raise AllModelsExhausted()Useful in production, not for cost optimization — you lose quality on fallback.
Pattern 5 — Task-Specific Routing
Skip generic complexity scoring when you know the task type:
| Task | Route | |------|-------| | Commit message from diff | Haiku | | Summarize 5-10 lines | Haiku | | Classify intent | Haiku | | Fix a failing test | Sonnet | | Write new feature | Sonnet | | Code review, architecture decision | Opus | | Multi-agent orchestration | Opus | | Complex debugging across systems | Opus |
Encode this as a map in code, not a prompt.
Anti-patterns
| Anti-pattern | Consequence | Fix | |--------------|-------------|-----| | Opus for everything | 10-40x bill | Start with Sonnet, measure, demote | | Haiku for code review | Misses subtle bugs | Sonnet minimum for code quality | | Router overhead > savings | Haiku classifier eats the margin | Skip router if >80% of traffic is one tier | | Different prompts per tier | Maintenance nightmare | Same prompt, just swap model | | No telemetry | Can't optimize | Log model + tokens + cost per request |
Measuring
Track per-route:
- Cost per request
- Latency p50/p95
- Quality score (human-labeled or auto-evaluated)
- Escalation rate (how often you fell back to a bigger model)
Target: move the Pareto curve — cheaper at equal quality OR better at equal cost.
Related
- `llm-ops-engineer` agent — production routing strategy
- `prompt-cach
Read more
name: model-routing-patterns description: "Multi-model pipelines (Haiku/Sonnet/Opus): cost routing, escalation, fallback chains. Triggers: model routing, Haiku, Sonnet, Opus, escalation, fallback chain." effort: medium user-invocable: false allowed-tools: Read
Model Routing Patterns
Three Claude tiers. Using Opus for everything is 10-40x more expensive than it needs to be. Using Haiku for everything loses accuracy on hard tasks. The craft is routing.
Model Characteristics (2026)
| Model | $/1M in·out | Cost (rel.) | Strengths | When | |-------|-------------|-------------|-----------|------| | Haiku 4.5 | $1 / $5 | 1x | Classification, extraction, simple tools, moderation | Bulk processing, triage, labels | | Sonnet 5 | $3 / $15 | ~3x | General coding, reasoning, most agent tasks | Default workhorse | | Opus 4.8 | $5 / $25 | ~5x | Complex reasoning, orchestration, architecture, large context | Hard, rare, high-stakes | | Fable 5 | $10 / $50 | ~10x | Most demanding long-horizon agentic work | Only when explicitly chosen |
Prices are per 1M tokens; ratios are approximate and shift between releases. Re-check pricing before committing a production path.
> **Fable 5 is not the default "best model".** Its price sits above Opus-tier, and Opus 4.8 is state-of-the-art on planning/orchestration at half the input and output cost. Reach for Fable 5 only when the user explicitly asks for it or a benchmarked task genuinely needs it — for "use the strongest model", the target is `claude-opus-4-8`.
Effort — the cheaper lever before swapping models
On Fable 5 / Opus 4.8 / Sonnet 5, `output_config.effort` (`low` | `medium` | `high` | `xhigh` | `max`) controls thinking depth and token spend **without changing the model** — so it does not invalidate the prompt cache the way a mid-session model swap does. Tune effort first; drop to a cheaper model only when effort alone can't hit the cost target.
| Effort | Use for | |--------|---------| | `low` | Latency-sensitive, non-intelligence-sensitive: chat, simple lookups, cheap subagents | | `medium` | Cost-conscious step-down from the default | | `high` | Default for most intelligence-sensitive work (a good quality/cost balance) | | `xhigh` | Hardest coding and agentic tasks (Claude Code's default) | | `max` | Correctness matters more than cost; test for diminishing returns |
In our agents, effort is set per skill/agent frontmatter (`effort:`), not swapped at runtime. Combine effort routing with model routing: e.g. `sonnet` at `high` often beats `opus` at `low` for cost-equal quality — benchmark before committing.
Pattern 1 — Complexity Router (pre-classify)
Cheap model classifies the request, then routes to the right tier:
def route(user_message: str) -> str:
complexity = classify_with_haiku(user_message) # returns: simple | medium | hard
return {"simple": "haiku", "medium": "sonnet", "hard": "opus"}[complexity]Good when ~60% of traffic is simple. Overhead: one Haiku call per request (~100 tokens).
Pattern 2 — Confidence-Based Escalation
Try the cheap model first, escalate only when it hesitates:
def solve(problem: str):
haiku = call_haiku(problem)
if haiku.confidence > 0.85:
return haiku.answer
sonnet = call_sonnet(problem + haiku.reasoning)
if sonnet.confidence > 0.8:
return sonnet.answer
return call_opus(problem)Haiku must be prompted to output confidence (e.g. via tool-use structured output — see `json-mode-patterns`). Pure self-reported confidence is noisy; combine with a heuristic (output length, tool calls, hedging words).
Pattern 3 — Sub-agent Delegation (Opus orchestrates, Haiku workers)
Orchestrator reasons about the plan, workers execute atomic steps:
Opus (planner) ├── Haiku (extract_dates_from_doc_1) ├── Haiku (extract_dates_from_doc_2) ├── Haiku (extract_dates_from_doc_3) └── Opus (synthesize all extractions into timeline)
Real example: `/orchestrate` in ai-toolkit runs Opus as planner, subagents (model per agent's frontmatter) as workers. See `app/agents/*.md` — each agent sets `model:` explicitly.
Pattern 4 — Fallback Chain (resilience, not cost)
When primary is rate-limited or errors, degrade gracefully:
def call_with_fallback(messages):
for model in ["claude-opus-4-8", "claude-sonnet-5", "claude-haiku-4-5"]:
try:
return client.messages.create(model=model, messages=messages, ...)
except (RateLimitError, OverloadedError):
continue
raise AllModelsExhausted()Useful in production, not for cost optimization — you lose quality on fallback.
Pattern 5 — Task-Specific Routing
Skip generic complexity scoring when you know the task type:
| Task | Route | |------|-------| | Commit message from diff | Haiku | | Summarize 5-10 lines | Haiku | | Classify intent | Haiku | | Fix a failing test | Sonnet | | Write new feature | Sonnet | | Code review, architecture decision | Opus | | Multi-agent orchestration | Opus | | Complex debugging across systems | Opus |
Encode this as a map in code, not a prompt.
Anti-patterns
| Anti-pattern | Consequence | Fix | |--------------|-------------|-----| | Opus for everything | 10-40x bill | Start with Sonnet, measure, demote | | Haiku for code review | Misses subtle bugs | Sonnet minimum for code quality | | Router overhead > savings | Haiku classifier eats the margin | Skip router if >80% of traffic is one tier | | Different prompts per tier | Maintenance nightmare | Same prompt, just swap model | | No telemetry | Can't optimize | Log model + tokens + cost per request |
Measuring
Track per-route:
- Cost per request
- Latency p50/p95
- Quality score (human-labeled or auto-evaluated)
- Escalation rate (how often you fell back to a bigger model)
Target: move the Pareto curve — cheaper at equal quality OR better at equal cost.
Related
- `llm-ops-engineer` agent — production routing strategy
- `prompt-cach
Professional-grade AI coding toolkit with multi-platform support. Machine-enforced safety, 109 skills, 44 agents, expanded lifecycle hooks, persona presets, experimental opt-in plugin packs, and benchmark tooling — works with Claude Code, Claude Chat/Cowork,
Repo: softspark/ai-toolkit
Other skills on ai-toolkit.
- /ai-toolkit-rules
Mandatory engineering, security, testing, git, performance, quality, and response rules. Claude MUST load this skill for every technical, coding, debugging, review, architecture, DevOps, data, or file-editing task in Chat or Cowork.
Open skill - /mem-search
Search past coding sessions using natural language. Finds relevant observations, decisions, and context from previous work.
Open skill - /a11y-validate
Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. Triggers: a11y, accessibility, WCAG, EAA, ARIA, contrast, keyboard, screen reader.
Open skill - /agent-creator
Creates new specialized agents with frontmatter, tools, delegation. Triggers: new agent, create agent, agent scaffold, specialized agent.
Open skill - /analyze
Analyzes code quality, complexity, patterns across codebase. Triggers: quality report, hotspot scan, code analysis, architecture signal.
Open skill - /api-patterns
REST/GraphQL API design: naming, versioning, pagination, idempotency, OpenAPI. Triggers: API design, REST, GraphQL, OpenAPI, Swagger, idempotency, rate limit.
Open skill

