PriceIndex

Token Pricing / Guide

AI token pricing explained: calculate the real cost of your workload

Input, output, caching and retries all shape your AI bill. A practical guide to comparing models using a consistent workload, with a transparent worked example.

PriceIndex EditorialPublished Oct 3, 2026Updated Oct 3, 20267 min read
An editorial illustration of a processor with green and yellow data streams

An editorial illustration of a processor with green and yellow data streams

The short answer

An AI API bill is usually a combination of input tokens and output tokens, multiplied by the rates for the model and billing mode you actually use. Input is what you send; output is what the model returns. A low input price does not guarantee the lowest total bill, especially when a task produces long answers.

The most useful comparison starts with a workload, not a leaderboard. Measure the typical prompt length, expected response length, request count and retry rate. Then compare the same workload across candidate models. This guide explains the calculation without claiming that any one provider is universally cheapest.

  • Separate input and output rates instead of comparing a single headline price.
  • Label cached input, batch discounts and tool charges separately.
  • Use a task-specific quality threshold before optimizing purely for cost.

Input, output and context are different

Input tokens can include your system instructions, user messages, conversation history, retrieved documents and tool results. In a multi-turn chat, the same earlier conversation may be sent repeatedly. A short new message therefore does not necessarily mean a short billed input.

Output tokens are the generated answer and, depending on the provider and model, may include billed reasoning tokens that are not displayed as ordinary answer text. Check the provider documentation and usage response rather than assuming visible words equal billed output.

A context window is a capacity limit, not a prepaid allowance. A model that accepts a large context does not automatically include that volume free of charge. Maximum output is another separate limit: a large input capacity does not imply an equally large possible response.

  • Token counts vary by tokenizer, language, formatting and modality.
  • Do not convert words to tokens using one universal fixed ratio.
  • Use API usage counters or the provider tokenizer for estimates.

The cost formula

For standard text-token billing, estimated cost = (uncached input tokens / 1,000,000 × input price per million) + (output tokens / 1,000,000 × output price per million). If cached input has a separate rate, split that volume into its own term rather than billing the same tokens twice.

Use a single currency and the same rate unit throughout the calculation. Some APIs display prices per thousand tokens; others use per million. Multiply a per-thousand rate by 1,000 before comparing it with a per-million rate. Taxes, currency conversion and platform charges are outside this simple token calculation.

The model cost calculator on PriceIndex applies your chosen input and output volumes to the current rates displayed in our catalog. A forecast remains an estimate: actual usage can differ because of retries, response length, routing, optional features or a later price change.

A worked example: 10,000 requests

Consider an illustrative workload of 10,000 requests per month. Each request sends 2,000 input tokens and produces 500 output tokens. That totals 20 million input tokens and 5 million output tokens. These volumes and rates are hypothetical teaching examples, not a quote for a named provider.

Suppose Model A charges $0.50 per million input tokens and $2.00 per million output tokens. Input costs $10 and output costs $10, giving an estimated token bill of $20. Suppose Model B instead charges $0.20 for input and $4.00 for output. Its input costs $4, but output costs $20, bringing the total to $24.

Model B has the cheaper input rate, yet it is more expensive for this workload. If the average answer becomes longer, the output rate matters even more. Conversely, a retrieval-heavy workload with long prompts and brief answers may be much more sensitive to the input rate.

Now assume 10% of requests require one additional full attempt with the same average token volumes. The forecast grows to 22 million input and 5.5 million output tokens. Model A rises to $22 and Model B to $26.40. Real retries may consume different token amounts; measure them rather than applying this percentage blindly.

  • Model A: 20 × $0.50 + 5 × $2.00 = $20.
  • Model B: 20 × $0.20 + 5 × $4.00 = $24.
  • Compare total cost for an identical workload, not just the cheapest column.

Caching and batch: savings with conditions

Prompt caching can reduce the cost of repeated eligible input. Eligibility can depend on prefix matching, minimum prompt size, cache lifetime and provider-specific behavior. Some services charge for cache writes or storage in addition to reads. A advertised cached-input rate is therefore not a discount on every token you send.

Batch processing is useful when responses are not needed immediately. Providers may price asynchronous requests differently from interactive requests, but discounts, completion windows and supported features vary. Do not apply a batch rate to a real-time application unless its requests actually qualify.

For a reliable forecast, record how much input really receives a cache hit and how much traffic can move into asynchronous processing. Keep standard, cached and batch rates as separate scenarios. This is more informative than assuming an ideal discount across the entire monthly bill.

  • Track cache-hit tokens, not just requests that reuse similar text.
  • Check write fees, retention rules and minimum eligible lengths.
  • Keep latency-sensitive traffic separate from batch jobs.

What the token total can leave out

Tool use may add charges that are not represented by an input/output token rate. Search requests, code execution, file storage and hosted retrieval can each have their own billing meters. A model may support a tool without including the tool service free of charge.

Images, audio and video can use modality-specific billing. Depending on the API, the bill may be expressed in tokens, images, seconds, minutes or another unit. Avoid comparing a per-image price directly with a text-token rate without a documented conversion for the exact request.

Your surrounding stack also contributes to the bill. A retrieval application can incur embedding charges, database compute, vector storage, object storage, bandwidth and monitoring. A self-hosted model replaces an API invoice with infrastructure and operational costs; it does not eliminate costs altogether.

  • Check extra meters on the official pricing page.
  • Include failed requests and retries according to actual billing behavior.
  • Budget for retrieval, storage and network transfer separately.

Compare quality and speed alongside price

A lower-cost model is useful only if it reliably completes your task. Evaluate representative examples from your application, including edge cases and your longest prompts. Track correctness, formatting failures, refusal behavior and how often a human or a stronger model must repair an answer.

Benchmarks help narrow a shortlist, but they are not interchangeable. A reasoning score, coding evaluation and preference-based arena rating measure different things. Check the benchmark definition and coverage before drawing conclusions. Missing scores should remain missing, not be treated as zero quality.

Latency also has several dimensions. Time to first token affects how quickly a streamed answer begins, while output speed affects how quickly it finishes. Long responses and tool round trips can dominate total wall-clock time. Measure the end-to-end user experience under your own traffic conditions.

  • Set a minimum task-success rate before comparing costs.
  • Measure cost per successful result, not only cost per attempt.
  • Use benchmark scores as evidence, not a universal verdict.

A practical production checklist

Start with a small representative trial and capture usage counters for every attempt. Group costs by model, endpoint, application feature and billing mode. This makes it easier to spot a long system prompt, retrieval payload or unusually verbose feature that silently dominates spend.

Set reasonable output limits and monitor repeated attempts. Test shorter prompts and smaller retrieval payloads against the same quality checks before rolling out changes. Routing easy tasks to a less expensive model can help, but an inaccurate router can create enough extra retries to erase the saving.

Finally, revisit your assumptions when rates change or a model is retired. Record which rates and dates were used in a forecast. Use the model pages, comparison workspace and price tracker to review catalog information, then confirm the applicable terms with your provider before making a purchasing decision.

  • Capture input, output, cached usage and retries.
  • Calculate a baseline and a higher-usage scenario.
  • Compare two or three models with the same workload.
  • Review price history and official billing terms before deployment.

Sources & further reading

The details

Frequently asked questions

Input tokens represent the content sent to a model, including instructions and context. Output tokens represent generated content; some providers also bill reasoning tokens. The exact token count depends on the tokenizer and API usage accounting.