Skip to content
OTFotf
All posts

Claude Haiku 5.5 production routing: measure cost, latency, and task quality

D
DaveAuthor
6 min read
Claude Haiku 5.5 production routing: measure cost, latency, and task quality

Claude Haiku 5.5 is a candidate for the parts of an application that run often, need a quick response, and have a bounded definition of success. Anthropic positions it for classification, extraction, routing, summaries, compaction, and subagent work. That is a useful starting list, not a reason to send every inexpensive request to it.

The practical decision is narrower: which calls in your product can move to Haiku 5.5 without pushing error rates, retries, or human review costs above the savings? Measure the whole task on your own inputs, including the larger token counts the new tokenizer can produce, then roll out behind a model setting with a fallback path.

Dex and Nova sorting bounded app tasks into a fast low-cost lane and an escalation lane

Start with work that has a narrow answer

Good first candidates have an output you can validate cheaply. Examples include assigning a support request to one of a fixed set of queues, extracting a handful of fields from a document, deciding which workflow should handle a request, or summarizing a known record for a user. A small model can also perform a bounded subtask while a larger model handles the main job.

Write down the contract before changing the model. For a classifier, that might be an allowed label set plus an abstain option. For extraction, define which fields may be null and what counts as a valid value. For routing, record the eligible destinations and the rule for sending an ambiguous request to a person or a stronger model. A vague prompt with a subjective “good answer” makes model comparisons hard to interpret.

Do not begin with an open-ended coding agent, a long reasoning chain, or an action that changes customer data. Anthropic says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding. Keep those workloads where they are until your own evaluation gives you a reason to split out a small, well-defined subtask.

Build a test set from real traffic

Sample representative requests from your existing workload. Include ordinary cases, common edge cases, malformed input, and examples where the correct answer is to abstain or escalate. Remove or protect personal data according to your retention rules. Freeze the expected result and evaluation method before comparing models so a new prompt does not quietly move the goalposts.

For each candidate request, record the outcome you care about. A routing task might require the correct destination and a safe abstention on uncertain cases. An extraction task might score field-level precision and recall, with stricter weighting for fields that trigger a downstream action. A summary might be checked for required facts, unsupported statements, and whether a user can complete the next step.

Run the current model and Haiku 5.5 on the same cases. If effort level or prompt wording changes, record each configuration separately. Haiku 5.5 supports adjustable effort; start with the level that fits the task and compare alternatives on the same frozen cases. A faster response that produces more retries or bad escalations may cost more per completed task than a slower first response.

Luna comparing evaluation cases, task success, response time, and cost before routing changes

11 production screens. Login, database, payments — all wired.

The SaaS Dashboard Kit ships everything already connected. Nothing to set up. Live demo at saas.otf-kit.dev.

See the live demo

Recompute the bill from tokens, not the old estimate

Haiku 5.5 uses a newer tokenizer. Anthropic says the same input text produces approximately 30% more tokens than it did with Haiku 4.5, though the exact increase depends on the content. Recount representative prompts with Haiku 5.5 rather than multiplying the old token count by a fixed factor. The per-request usage fields and token-counting results are the numbers to use for your estimate.

Pricing also changes at a 100,000-token prompt threshold. For prompts up to that size, the listed rates are $0.10 per million input tokens and $0.50 per million output tokens. Above the threshold, they are $0.50 and $2.50. Cache writes and reads have separate rates, and the Batch API offers a discount on input and output tokens. Keep each price tier and cache category separate in your spreadsheet; a single blended rate can hide long prompts that fall into the higher tier.

For example, if a short-prompt workload uses 40 million input tokens and 5 million output tokens in a month, the listed input and output charges are $4 and $2.50, or $6.50 before cache charges and any other applicable costs. Calculate the old model’s cost from its own recorded usage, then compare cost per successfully completed task. Do not apply Anthropic’s “around 75% less on average” statement to your traffic without measuring your prompt mix and token usage.

Track latency and failures alongside spend

Log the model ID, effort setting, input and output token counts, prompt-size tier, cache usage, response time, completion status, and whether the request retried or escalated. For user-facing work, inspect p50 and p95 latency separately; a good median can hide a slow tail that users notice. For background work, track completion time against the job’s deadline.

Use a cost-per-success measure as well as cost per call. One simple form is:

cost per successful task = total model cost / tasks that pass the task-specific quality check

Include retry and fallback calls in total model cost. If a classifier is cheaper per call but sends more requests to a human queue, include that added review work in the decision. Keep model output checks close to the action: validate an extracted enum before saving it, and require an explicit policy for low-confidence or malformed results.

Put the model choice behind configuration

Keep the model ID, effort, and fallback policy in application configuration rather than scattering them through request code. Anthropic’s Claude API model ID is claude-haiku-5-5; other platforms may use a different provider-specific identifier. Make the deployment target choose the appropriate ID, and log the resolved value so a console display name is never mistaken for the ID your application actually sent.

Start with shadow evaluation if the workflow allows it: send a copy of eligible requests to Haiku 5.5, but let the current model’s answer remain the one that reaches the user or changes state. Compare results on the frozen set and review disagreements. For interactive tasks, move a small percentage of eligible traffic to Haiku 5.5 and keep the existing route ready for rollback. For asynchronous tasks, migrate one queue or job type first.

Set a stop condition before rollout. Examples are a minimum task success rate, a maximum p95 latency, an allowed escalation rate, and a cost-per-success ceiling. If the task misses any threshold, route it back and inspect the cases that failed. Avoid changing the model, prompt, effort level, and validation logic in one release; otherwise it becomes difficult to tell which change moved the result.

Make the routing decision workload by workload

Haiku 5.5 is worth evaluating where calls are frequent, latency matters, and you can define what a correct result looks like. Keep the existing model for tasks whose quality is hard to grade or where a wrong answer has a large cost. If you already use Opus 5.5 for complex work, consider Haiku for a bounded subtask around that flow rather than replacing the lead model wholesale. Our Opus 5.5 cost and model-ID guide covers the adjacent model’s pricing and pinning choices.

The decision should come from a representative replay and a staged production measurement, not from the model name or a benchmark table. Recount tokens on the new tokenizer, keep the 100,000-token price boundary visible, and compare the cost of successful outcomes. That gives you a concrete answer to where a smaller model belongs in your application—and a rollback path if it does not earn that place.

Sources

ai-toolsagentsbackend
OTF SaaS Dashboard Kit

Ship the product, not the setup.

  • 11 production screens — auth, billing, team, analytics, settings
  • Real database, payments, and login — all wired on day 1
  • AI configs pre-tuned so your agent extends instead of regenerates
Need more than components?

Full-stack kits.
Pay once, own the code.

Auth, database, and payments already connected — so you ship product, not setup. Or take the delivered kits in the Bundle.

Everything Bundle — $149See full pricing

Get the free AI configs pack

Pre-tuned AI configs for Cursor, Claude, and Lovable — drop them in and your AI tool instantly understands your project.

No spam. Unsubscribe any time.

Prefer the free SDK? Star it on GitHub →