---
title: "Route by task, not by loyalty"
date: 2026-06-30
series: 15
summary: "The biggest cost lever in an agent stack right now is a router that knows when not to call the frontier model."
voice: architect
tags: [mid-tier-agent-economics, capability-routing, effort-dial-cost, cost-per-task]
image: "/notes/capability-routing-sonnet5/image.png"
imageAlt: "split panel: every task hitting the frontier model vs a two-tier router with effort caps, plus the launch pricing strip"
linkedin: null
sources:
  - title: "Sonnet 5 announcement, pricing ($2/$10 intro through Aug 31, then $3/$15), \"close to Opus 4.8 at lower prices\": Anthropic"
    url: https://www.anthropic.com/news/claude-sonnet-5
    date: 2026-06-30
  - title: "Launch coverage, agentic positioning and pricing analysis: TechCrunch"
    url: https://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/
    date: 2026-06-30
  - title: "Benchmark comparison vs Sonnet 4.6 and Opus 4.8 (Terminal-Bench 2.1: 80.4 vs 67.0; agentic coding 63.2 vs 69.2 Opus): MarkTechPost"
    url: https://www.marktechpost.com/2026/06/30/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared/
    date: 2026-06-30
dateApprox: false
---

The biggest cost lever in an agent stack right now is a router that knows when not to call the frontier model.

Anthropic shipped Sonnet 5 last week at $2 in / $10 out per million tokens intro pricing ($3/$15 from September). Vendor numbers put it just under Opus 4.8 on agentic coding and ahead on some knowledge work. Terminal-Bench went from 67.0 on Sonnet 4.6 to 80.4.

Mid-tier models now clear the bar for most production agent loops. That changes the architecture more than the bill. What I would do with it:

1. **Route by task, not by loyalty.** Tool-call planning, extraction, classification go mid-tier. Deep multi-file reasoning escalates. Two tiers beat a monolith on cost per task.

2. **Meter cost per task, not per token.** Token price tells you little once models think for variable lengths. Tag every run with its task class and its dollars. Alert on outliers.

3. **Re-run your evals before flipping traffic.** The gap to frontier is benchmark-thin, and your workload is not a benchmark. Golden set first, then migrate.

We run the same tiering on-prem with open weights. Same math, different price tags.

**The trap:** the effort dial is a budget, not a quality knob. Uncapped, the cheap model on a hard task can spend past Opus and hand you the same answer.

Cost per task is the only price that matters in an agent stack.
