What is GPT-6 Luna Token Counter?
This GPT-6 Luna token counter page explains how to measure input for gpt-6-luna, OpenAI's model for focused, high-volume tasks. Prompt length matters when a short task runs thousands of times: input, output and cache categories all contribute to API cost.
OpenAI lists a 1,050,000-token context window and 128,000 maximum output tokens for Luna. The available context is a model limit, not a promise that a word count or another model's tokenizer can measure your complete Luna request.
GPT-6 Luna Tokenizer Support in Tiktokenizer
Support level: reference-only. Tiktokenizer's existing runtime does not contain a verified GPT-6 Luna tokenizer mapping. The local explorer uses o200k_base for text inspection; its token IDs belong to that encoding and are not presented as Luna token IDs.
Exact means a verified tokenizer is counted directly. Compatible means an established base-encoding mapping is used with stated limits. Luna is not assigned either label here. Use OpenAI's Responses input-token count with model gpt-6-luna for the same structured input your application will send.
GPT-6 Luna Pricing and Cost Notes
Facts checked on 2026-10-08. The table lists USD per 1M tokens for Standard processing with up to 272,000 input tokens, from the official GPT-6 Luna model document.
Above 272,000 input tokens, the full request uses 2x input and cache rates and 1.5x output rates. Processing tiers and regional choices can also change prices. Apply a cached-input price only to tokens actually classified as cached by OpenAI; pasting the same text twice in this editor does not establish an API cache hit.
| Usage category | USD per 1M tokens |
|---|---|
| Input | $0.10 |
| Cache read | $0.01 |
| Cache write | $0.125 |
| Output | $0.50 |
How to Count GPT-6 Luna Tokens
- 1. Paste the prompt in the reference text explorer and inspect its live raw-encoding count and token IDs.
- 2. Switch encodings if you want to compare splits; do not rename a reference result as an exact Luna count.
- 3. In your own application, count the complete request using OpenAI's Responses input-token counting operation with model gpt-6-luna, including supported messages, tools and media.
- 4. Reserve room for output and reasoning, then check the actual response usage before applying Luna's input, output and cache prices. This browser does not send the prompt to OpenAI.
Token Count vs Billed Usage
A local text encoding omits Luna's model-specific chat template and request serialization. System instructions, prior turns, tool definitions and images can change the model's input count. Use the complete structured payload for an authoritative preflight input count.
Generated output and internal reasoning cannot be known from the prompt alone. Cache reads and writes depend on the request and provider cache state; tool calls may carry additional charges. For a cost estimate, keep these usage categories separate and use the actual model and processing tier. A ChatGPT free plan does not make Luna API calls free.
Tiktokenizer Reference Sources
Tokenizer behavior and model limits change. Verify production decisions with current provider documentation: