Opus gets the launch-day headlines. Sonnet is the model that actually runs your outbound — $2 per million input tokens, $10 output, a 1M-token context window, and "fast" latency. At that price-to-intelligence ratio, it's the default tier for production automation.
Shipped Sept 28, 2026 · Knowledge cutoff June 2026
At $2/$10, Sonnet 5.5 is half the cost of Opus 5.5 and a fraction of Fable — with far more intelligence and a 5× larger context than Haiku. For the tens of thousands of small AI calls outbound actually runs on, this is the tier the math works at.
Base API pricing per million tokens. Batch runs 50% cheaper; prompt-cache reads cost 10% of base input — both compound once you process thousands of leads a day.
Outbound isn't one big AI call. It's tens of thousands of small ones: a personalization line per prospect, an enrichment judgment per Clay row, a reply classified, a transcript summarized, a next step decided. At that volume, the model's price per token is your margin.
The Opus tier is overkill for most of that work, and its price shows it. Haiku is cheap but you feel the intelligence gap on anything requiring real reasoning about a prospect. Sonnet 5.5 is the tier built for the middle: enough judgment to write and decide well, fast enough to keep a campaign moving, cheap enough to run on every record instead of a sample.
Our read: this is the default model for production outbound automation. Reserve Opus 5.5 for the few steps that genuinely need deeper reasoning, and use Sonnet 5.5 for everything that runs on every lead. The right build usually mixes both tiers rather than paying Opus rates for work Sonnet handles fine.
Forced tool use (tool_choice set to any or tool) now returns a 400, the old disabled-thinking mode was replaced with a between-tools mode at high effort or below, and the earlier computer_20251124 tool isn't accepted. Nothing dramatic — but run a sample batch before you flip production traffic.
A first line or custom P.S. per prospect doesn't need Opus-level reasoning — it needs a model that reads the research and writes something human, across thousands of rows, without blowing the budget. At $2/$10 with batch pricing, per-lead personalization is economical at real volume. Core to how we build cold email infrastructure that doesn't read like a template.
Enrichment is where token costs quietly pile up — a model call on every row of a large table. Sonnet 5.5's 1M context lets you feed a full company page, a job posting, and prior notes into one judgment, and its speed keeps a 50,000-row table from taking all night. The enrichment brain behind serious list building. (See our take on Claygent Skills.)
Voice agents live and die on latency — a model that thinks for three seconds before answering kills a call. Sonnet 5.5's "fast" tier plus adaptive thinking is the profile you want: quick enough to hold a natural conversation, smart enough to handle an objection. A strong fit for the LLM layer inside the AI calling systems we run on Bland and Retell.
After the send and the call comes the sorting: is this reply a yes, a not-now, or an unsubscribe? Did that call book a meeting or need a human? Classification and summarization at inbox volume is exactly the work Sonnet 5.5 is priced for. Route interested replies to a rep, auto-handle the rest, and feed clean outcomes back into the CRM.
Cheap enough to personalize, enrich, and classify on every lead — not a sample. That's the difference between a demo and a campaign.
Sonnet 5.5 for the volume work, Opus 5.5 for the few steps that need deeper reasoning. Paying Opus rates for Sonnet-grade work is the quiet budget leak.
Batch requests run 50% cheaper and cache reads cost a tenth of input. A pipeline built to use both pays far less than the sticker rate.
tool_choice values any and tool now return a 400, disabled-thinking mode changed, and the older computer-use tool isn't accepted. Run a sample batch, confirm output quality and tool behavior, then move production traffic.We build and run outbound systems for B2B teams — cold email infrastructure, AI calling on Bland and Retell, and Clay-based list building. Picking the right model for each step, and wiring it so the cost math works at volume, is the everyday job.