GPT-6 Luna · reference-only

GPT-6 Luna Token Counter

Plan GPT-6 Luna input size and understand its API prices. Inspect text locally with a reference encoding, then confirm the model-specific count with OpenAI.

Accuracy note

reference-only for GPT-6 Luna. The selected raw encoding is not verified as Luna's tokenizer. Use OpenAI's model-aware input count and final API usage for request and billing decisions.

Official counting docs →
API model ID
gpt-6-luna
Context
1,050,000 tokens
Maximum output
128,000 tokens
Local support
reference-only

Facts checked against official sources on .

What is GPT-6 Luna Token Counter?

This GPT-6 Luna token counter page explains how to measure input for gpt-6-luna, OpenAI's model for focused, high-volume tasks. Prompt length matters when a short task runs thousands of times: input, output and cache categories all contribute to API cost.

OpenAI lists a 1,050,000-token context window and 128,000 maximum output tokens for Luna. The available context is a model limit, not a promise that a word count or another model's tokenizer can measure your complete Luna request.

GPT-6 Luna Tokenizer Support in Tiktokenizer

Support level: reference-only. Tiktokenizer's existing runtime does not contain a verified GPT-6 Luna tokenizer mapping. The local explorer uses o200k_base for text inspection; its token IDs belong to that encoding and are not presented as Luna token IDs.

Exact means a verified tokenizer is counted directly. Compatible means an established base-encoding mapping is used with stated limits. Luna is not assigned either label here. Use OpenAI's Responses input-token count with model gpt-6-luna for the same structured input your application will send.

GPT-6 Luna Pricing and Cost Notes

Facts checked on 2026-10-08. The table lists USD per 1M tokens for Standard processing with up to 272,000 input tokens, from the official GPT-6 Luna model document.

Above 272,000 input tokens, the full request uses 2x input and cache rates and 1.5x output rates. Processing tiers and regional choices can also change prices. Apply a cached-input price only to tokens actually classified as cached by OpenAI; pasting the same text twice in this editor does not establish an API cache hit.

Usage categoryUSD per 1M tokens
Input$0.10
Cache read$0.01
Cache write$0.125
Output$0.50

How to Count GPT-6 Luna Tokens

  • 1. Paste the prompt in the reference text explorer and inspect its live raw-encoding count and token IDs.
  • 2. Switch encodings if you want to compare splits; do not rename a reference result as an exact Luna count.
  • 3. In your own application, count the complete request using OpenAI's Responses input-token counting operation with model gpt-6-luna, including supported messages, tools and media.
  • 4. Reserve room for output and reasoning, then check the actual response usage before applying Luna's input, output and cache prices. This browser does not send the prompt to OpenAI.

Token Count vs Billed Usage

A local text encoding omits Luna's model-specific chat template and request serialization. System instructions, prior turns, tool definitions and images can change the model's input count. Use the complete structured payload for an authoritative preflight input count.

Generated output and internal reasoning cannot be known from the prompt alone. Cache reads and writes depend on the request and provider cache state; tool calls may carry additional charges. For a cost estimate, keep these usage categories separate and use the actual model and processing tier. A ChatGPT free plan does not make Luna API calls free.

Tiktokenizer Reference Sources

Tokenizer behavior and model limits change. Verify production decisions with current provider documentation:

Frequently asked questions

Can Tiktokenizer count GPT-6 Luna exactly in the browser?

No. Local Luna support is reference-only until its tokenizer mapping is verified. The explorer's count and IDs belong to the selected raw encoding.

How much does GPT-6 Luna cost per million tokens?

As checked on 2026-10-08, Standard prices up to 272,000 input tokens are $0.10 input, $0.01 cache read, $0.125 cache write and $0.50 output per 1M tokens. Longer requests use higher rates.

Can I reuse a GPT-5.6 Luna count?

Treat it as a comparison only. Count the actual GPT-6 Luna structured request with OpenAI instead of assuming a previous model's local configuration matches.

Why can Luna API usage exceed my pasted-text count?

The request may include chat formatting, system instructions, tools, images and other input outside the editor. Generation also adds output and reasoning usage.