How the token cost calculator works
Paste text to count it locally with the site's existing tokenizer, or switch to manual input tokens for measured API usage. Enter estimated output tokens and calls per month independently. A call count alone cannot determine cost without a per-call input and output budget.
The live results update as you edit. Choose a model to see its per-call and monthly cost, then compare the same token quantities across all five rows. Use GPT-6 Sol, GPT-6 Luna and GPT-5.6 as examples of the same API budgeting workflow; this tool does not depend on a launch announcement.
Cost per request = (ordinary input × input rate + cached input × read or write rate + output × output rate) / 1,000,000. Monthly cost = cost per request × monthly calls. Monetary arithmetic uses integer currency units and preserves small token charges instead of rounding each category.
Example results: API cost per request and per month
The example text below is 37 raw o200k_base tokens. The default budget uses 1,000 output tokens and 10,000 monthly calls, with caching off. The comparison holds token quantities constant; equal text can require a different provider token count.
Facts checked on 2026-10-09. Costs are USD, computed from each linked model's rates. The Batch columns apply the documented discount. This worked example, the rate table and FAQ are included in the initial HTML.
Summarize this customer request in three bullet points, identify the next action, and return a short JSON response. Customer: I need to change my delivery address before the order ships.API pricing by model and input tier
Facts checked on 2026-10-09. All prices below are USD per 1M tokens for Standard first-party API processing. Each price links to its official model document. The GPT-5.6 Sol rate is the currently documented promotional rate; recheck the source before forecasting a future period.
For these models, total input up to and including 272,000 tokens uses the short-context schedule. Above that boundary the entire request uses the long-context schedule, including cached input and output. A large output allowance does not select the input tier. Batch uses half the applicable rates; regional processing and separately billed tools are outside this calculation.
Cache reads, cache writes and Batch discounts
Enable caching, choose read or write, and enter the prefix within your total input. A cached token is charged once in its cache category; it is not charged again as ordinary input. Read mode assumes a matching hit. Write mode prices the selected prefix at the write rate. Cache misses use ordinary input pricing.
Facts checked on 2026-10-09. The minimum eligible prefix for GPT-5.6 and later is 1,024 visible tokens. A shorter prefix uses ordinary input in this calculator. The default prefix field proposes 80% of your input, rounded down; edit it to match your actual usage. A local text count cannot establish cache eligibility, a hit or the correct billing tier.
The Batch toggle applies the published 50% discount for asynchronous API work. Creating a large number of synchronous calls does not qualify for that discount. Monthly totals repeat the chosen per-call scenario: an all-hit estimate excludes initial and renewed cache writes. Calculate writes and reads separately for a mixed workload rather than assuming one cache lasts the whole month.
| Cost assumption | Official reference |
|---|---|
| Cache categories and minimum prefix | OpenAI prompt caching · checked 2026-10-09 |
| Batch discount and asynchronous completion | OpenAI Batch API · checked 2026-10-09 |
Three hand-calculated cost checks
These independent expected totals are also asserted by the calculation tests. They cover Standard pricing, a cached read with Batch pricing, and a cache write above the input threshold with Batch pricing. No category or per-call subtotal is rounded before calculating a monthly total.
| Known inputs and model | Manual cost / call | Calls / month | Manual cost / month |
|---|---|---|---|
| GPT-6 Sol · 10,000 input + 1,000 output · no cache · Standard | $0.03 USD · checked 2026-10-09 · OpenAI rates | 10,000 | $300.00 USD · checked 2026-10-09 · OpenAI rates |
| GPT-6 Luna · 10,000 input (6,000 cache reads) + 2,000 output · Batch | $0.00073 USD · checked 2026-10-09 · OpenAI rates | 20,000 | $14.60 USD · checked 2026-10-09 · OpenAI rates |
| GPT-5.6 Terra · 300,000 input (200,000 cache writes) + 2,000 output · Batch | $0.718 USD · checked 2026-10-09 · OpenAI rates | 500 | $359.00 USD · checked 2026-10-09 · OpenAI rates |
Local token counts and billed API usage
The local text panel can switch between o200k_base and cl100k_base. Each raw encoding counts its supplied text exactly. GPT-6 model estimates remain reference-only, while this site's GPT-5.6 o200k_base view retains its compatible label. Choosing another comparison encoding does not establish model compatibility.
For a model-aware budget, count the complete structured input with the provider, then enter that count manually. Include history, system instructions and supported tools or media. Estimated output should include billed reasoning where applicable. Pasted plain text does not reveal hidden request formatting, image processing or future generated output.
After running the workload, replace the planning quantities with actual usage. Compare the model, service tier, cache categories and output on the same basis. This arithmetic estimates token charges; it does not validate a model's request-size limits or include subscription fees, taxes or separately metered tools.
Tiktokenizer Reference Sources
Tokenizer behavior and model limits change. Verify production decisions with current provider documentation:
- OpenAI GPT-6 Sol: model rates and long-context rules · checked 2026-10-09
- OpenAI GPT-6 Luna: model rates and long-context rules · checked 2026-10-09
- OpenAI GPT-5.6 Sol: model rates and long-context rules · checked 2026-10-09
- OpenAI GPT-5.6 Terra: model rates and long-context rules · checked 2026-10-09
- OpenAI GPT-5.6 Luna: model rates and long-context rules · checked 2026-10-09
- OpenAI pricing: current service tiers · checked 2026-10-09
- OpenAI prompt caching: prefix and billing categories · checked 2026-10-09
- OpenAI Batch API: discounted asynchronous requests · checked 2026-10-09
- OpenAI input token counting: complete request measurements · checked 2026-10-09