aeo-optimization
AI Engine Optimization - semantic triples, page templates, content clusters for AI citations
9-tier model routing system with cascading classifier fallback and result auto-evaluation
$ npx -y skills add alinaqi/claude-bootstrap --skill model-routing --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/model-routingContext preview
The summary Claude sees to decide when to auto-load this skill.
9-tier model routing system with cascading classifier fallback and result auto-evaluation
name: model-routing description: 9-tier model routing system with cascading classifier fallback and result auto-evaluation when-to-use: When configuring or debugging how prompts are classified and routed to the cheapest capable model user-invocable: false effort: medium
Every user prompt goes through a 9-tier classification pipeline before any AI model processes it. The system answers three questions:
1. **Which model should handle this?** — 9-tier cost/complexity classification 2. **Is the classifier itself working?** — Cascading fallback (qwen3 → kimi → deepseek → cache) 3. **Can we verify the result?** — Tool-level fallback + auto-evaluation
User types prompt
↓
UserPromptSubmit hook fires (~/.claude/hooks/route-task-hook)
↓
Classifier: qwen3 (local, free) classifies into tier
↓ (fails?)
Classifier: kimi (local, free) retries
↓ (fails?)
Classifier: deepseek-flash (~$0.0001) retries
↓ (fails?)
Classifier: cached tier from last success
↓
Hook injects routing decision into Claude's context
↓
Claude delegates to the right model or handles directly| Tier | Model | Input (per M) | Output (per M) | Handles | |------|-------|---------------|----------------|---------| | 0 | **Qwen3** (local) | $0 | $0 | grep, find, shell, syntax, log reading | | 1 | **Gemini 2.5 Flash-Lite** | $0.10 | $0.40 | Bulk extraction, classification, CIG pipelines | | 2 | **DeepSeek V4 Flash** | $0.14 | $0.28 | Simple code, CRUD, test writing, small fixes | | 3 | **DeepSeek V4 Pro** | $0.44 | $0.87 | Multi-file features, refactors, debugging (~80% of work) | | 4 | **Gemini 2.5 Flash** | $0.15 | $0.60 | Multimodal (images, video, audio), brand analysis | | 5 | **Kimi K2.6** | $0.60 | $2.50 | Code review, commit messages, diff summaries | | 6 | **Gemini 3.1 Pro + Search** | $1.25 | $10.00 | Deep research, Google grounding, 2M context | | 7 | **Codex** | varies | varies | Bulk generation, code review | | 8 | **Claude Sonnet/Opus** | $3-5 | $15-25 | Architecture, security, quality-critical |
When the hook says "delegate to X", run the matching command and return its output:
# Tier 0 — Qwen3 ~/bin/qwen3 "prompt" # Tier 1 — Gemini Flash-Lite ~/bin/gemini --flash-lite "prompt" # Tier 2 — DeepSeek Flash ~/bin/deepseek --flash "prompt" # Tier 3 — DeepSeek Pro ~/bin/deepseek --pro "prompt" # Tier 4 — Gemini Flash ~/bin/gemini --flash "prompt" # Tier 5 — Kimi ~/bin/kimi --quiet -p "prompt" # Tier 6 — Gemini Pro Search ~/bin/gemini --pro-search "prompt" # Tier 7 — Codex codex exec "prompt" # Tier 8 — Claude # Handle directly (no delegation)
Every `~/bin/` script follows the same pattern:
1. **Accepts prompt as argument**: `script "what is 2+2"` 2. **Model flags**: `--flash`, `--pro`, `--flash-lite`, `--pro-search` 3. **Quiet mode**: `--quiet` (where applicable) 4. **Output**: writes response to stdout, errors to stderr 5. **Exit codes**: 0 on success, non-zero on failure
~/bin/ ├── qwen3 # Shell: curl to local Ollama API ├── kimi # Shell: execs Kimi CLI binary ├── deepseek # Python: httpx to DeepSeek Anthropic-compat API ├── gemini # Python: httpx to Gemini OpenAI-compat API ├── research # Python: multi-backend research with auto-evaluation └── route-task # Shell: qwen3-powered task classification
The classifier itself can fail. When it does, cascading fallback kicks in:
| Level | Classifier | Cost | Threshold | |-------|-----------|------|-----------| | 1 | **qwen3** (Ollama) | $0 | 2s connect, 8s classify | | 2 | **kimi** CLI | $0 | Local process | | 3 | **deepseek-flash** | ~$0.0001 | API call | | 4 | **Cached tier** | $0 | From `~/.claude/routing-cache.json` |
The cache (`~/.claude/routing-cache.json`) saves the last successful tier and timestamp. After compaction, when Ollama may be briefly unreachable, the cache ensures routing continues without dropping to CLAUDE by default.
When Claude's built-in tools fail, external backends take over:
| Failed Tool | Fallback 1 | Fallback 2 | |-------------|------------|------------| | **WebSearch** / **WebFetch** | `~/bin/research "query"` | `~/bin/deepseek --pro "query"` | | **Read** / file access | `cat` via Bash | — | | **Grep** | `grep -r` via Bash | — |
Multi-backend research with auto-evaluation:
Maggy's `model_router.py` mirrors the same 9-tier structure in `DEFAULT_TIERS`. The `PiAdapter` uses the same delegation scripts for execution. Task type overrides in `routing_rules_defaults.py` ensure:
# Required for delegation scripts (in ~/.zshrc) export DEEPSEEK_API_KEY="sk-..." export GEMINI_API_KEY="..." # For gemini delegator export OPENAI_API_KEY="sk-..." # For codex CLI # Ollama must be running locally for qwen3 ollama serve # or launch at startup
Turn Claude Code into a self-reviewing, test-enforced engineering system that remembers context across sessions — then route work across 13 models from a single dashboard.
Repo: alinaqi/claude-bootstrap
AI Engine Optimization - semantic triples, page templates, content clusters for AI citations
Claude Code Agent Teams - default team-based development with strict TDD pipeline enforcement
Build AI agents with Pydantic AI (Python) and Claude SDK (Node.js)
Latest AI models reference - Claude, OpenAI, Gemini, Eleven Labs, Replicate
Android Kotlin development with Coroutines, Jetpack Compose, Hilt, and MockK testing