Per-request costs · Haiku 5.5 vs Haiku 4.5

Claude Haiku 5.5 Pricing Calculator

Enter prompt and output tokens to compare request costs, switch cache scenarios, and explore how a tokenizer change can affect your budget.

Accuracy note

Cost arithmetic uses the verified rates and the token quantities you enter. Both local Haiku text estimates are reference-only; cl100k_base is not a Claude tokenizer. Check the intended model with Anthropic count_tokens and reconcile costs with actual Messages usage.

Official counting docs →

Facts checked against official sources on .

Haiku 5.5 tiered pricing calculator

The live calculator starts with 100,000 total prompt tokens, 10,000 output tokens and an 80,000-token cacheable prefix. Edit the quantities to update the cost immediately; no Calculate button is required. The worked example below is also included in the initial HTML.

Use total prompt tokens, including cached input. At exactly 100,000 input tokens Haiku 5.5 uses the lower schedule. At 100,001 it uses the higher schedule for all input categories and output, rather than charging only the extra token at a higher rate. The comparison holds token quantities constant across models.

Example result: Haiku 5.5 vs Haiku 4.5 pricing

Example result for one request: 100,000 prompt tokens and 10,000 output tokens; cache scenarios allocate 80,000 input tokens to the cache and 20,000 to ordinary input. Cache writes use the 5-minute duration. The fixed example remains readable without JavaScript; live results use your edits.

Facts checked on 2026-10-09. Each cost below links to the official rate source. This compares equal token quantities, not equal text: actual Haiku 5.5 token counts can differ.

Haiku 5.5 pricing tiers and Haiku 4.5 pricing

Facts checked on 2026-10-09. Rates are USD per 1M tokens for standard first-party Claude API usage. Every price links directly to Anthropic's official pricing table. Batch discounts, regional modifiers and separately metered tools are outside this calculator.

The Haiku 4.5 column is the comparison baseline. Select the Haiku 5.5 column using total input length; output tokens do not select the tier. Recheck the official source when using this saved rate snapshot.

How prompt caching changes the cost

Choose no cache, cache write or cache read. The cacheable-prefix field is a subset of total prompt tokens, not extra input. A write charges that prefix once at the selected 5-minute or 1-hour write rate; a hit uses the read rate. The rest of the prompt uses ordinary input pricing. A read assumes an existing matching cache entry and excludes its earlier write cost.

Facts checked on 2026-10-09. Haiku 5.5 requires at least 512 tokens in the cacheable prefix; Haiku 4.5 requires 4,096. Below a model's minimum the calculator falls back to ordinary input charges for that model, following the official caching guide. A cache miss also uses ordinary input charges.

Total cost = uncached input × input rate + cache prefix × write or read rate + output × output rate, divided by 1,000,000. Each input token belongs to one category. With caching, the API input_tokens field alone omits cache reads and writes; add all three input usage fields before selecting the price tier.

Caching ruleOfficial documentation
Minimum prefix and cache-hit behaviorAnthropic prompt caching · checked 2026-10-09
Input, write and read usage fieldsAnthropic usage breakdown · checked 2026-10-09

Haiku 4.5 vs Haiku 5.5 tokenizer estimates

Paste the same text into the comparison below. This browser does not contain either official Haiku tokenizer: it uses cl100k_base as a proxy baseline, then rounds that baseline × 1.30 up to a whole token for Haiku 5.5. Both model rows are estimate / reference-only. The raw encoding row is exact for cl100k_base text only.

Facts checked on 2026-10-09. Anthropic's Haiku migration guide describes roughly 30% more tokens for the same text, with content-dependent variation. This illustrative multiplier is not a measured Haiku tokenizer difference and must not determine the 100k billing boundary. Use count_tokens separately for each model, or actual usage, for a model-aware comparison.

Example viewTokensSupport / method
Raw cl100k_base37exact for this encoding only
Haiku 4.5 proxy37estimate / reference-only: raw reference count
Haiku 5.5 illustration49estimate / reference-only: ceil(reference × 1.30)
Illustrated difference+12 (32.43%)Rounding changes the displayed percentage; actual model counts can differ
textExample text for both model estimates
Summarize this customer request in three bullet points, identify the next action, and return a short JSON response. Customer: I need to change my delivery address before the order ships.

Token counts, price tiers and billed usage

  • 1. Count the complete supported input separately with claude-haiku-4-5 and claude-haiku-5-5. Include system text, history, supported tools and media in the provider request.
  • 2. Enter total input and planned output in this calculator. With caching, total input is input_tokens + cache_creation_input_tokens + cache_read_input_tokens. Do not add cached input twice.
  • 3. Reserve output separately: visible answer length does not predict reasoning or final billed output. This calculator does not infer missing chat formatting, images or tool usage from pasted text.
  • 4. After the request, reconcile token categories with Messages usage and the official rates. The local text proxy is for exploring assumptions; it is not an invoice or a Claude tokenizer.

Tiktokenizer Reference Sources

Tokenizer behavior and model limits change. Verify production decisions with current provider documentation:

Frequently asked questions

Does Haiku 5.5 charge the higher rate only above 100,000 tokens?

No. The total input length selects the request's rate schedule. Exactly 100,000 uses the lower schedule; a longer prompt uses the higher schedule for ordinary input, output and cache categories. Cached input still contributes to total input length.

How does this compare Haiku 4.5 pricing with Haiku 5.5?

The cost calculator uses the same input, output and cache-prefix quantities for both models. The separate text module illustrates a possible count change; it does not silently replace the cost inputs. Equal-text costs require separately measured model counts.

Can I count the entire prompt at both input and cache rates?

No. A token is ordinary input, a cache write or a cache read within a request. Enter the total prompt and the prefix subset. Cache-read results assume a hit; prior creation is a separate request cost. Prefixes below a model's cache minimum use ordinary input charges.

Are the Haiku 4.5 and Haiku 5.5 text counts exact?

Both local Haiku views are estimate / reference-only. The baseline uses cl100k_base, which is not a Claude tokenizer; the Haiku 5.5 illustration applies a content-dependent approximate multiplier. Only the raw reference-encoding count is exact for that encoding. Use Anthropic count_tokens for model-aware preflight estimates and actual Messages usage for billing.

Will the calculator and example work without JavaScript?

The explanation, complete rate table, example request costs, tokenizer example and FAQs are pre-rendered in the initial HTML. Editing quantities, switching cache scenarios and pasting text require JavaScript. Prices are a dated snapshot with direct official source links.