A plausible monthly number can still rest on the wrong denominator.
The incident-response team expects 500 summaries a day. A spreadsheet assumes 2,000 input tokens and 500 output tokens per request, then applies one rate to the total. That forecast may be off because input and output can have different rates, and because the prompt may grow when incident notes or tool results are added. The arithmetic is only as sound as its counted tokens and rate units.
The example rates in this lesson are editable placeholders, not a quote from a vendor. We will assume $2 per million input tokens and $8 per million output tokens for the arithmetic. Replace them with the current rates for the exact model and billing path before using a real forecast.
- Known
- Expected request count, a draft prompt, and a target response format.
- Unknown
- Exact tokenizer counts, variable retrieved context, retries, tool calls, and actual output lengths.
- Decision
- Is an order-of-magnitude forecast enough for a pilot, and what must be measured before scaling?
A character, a byte, a token, and a context slot answer different questions.
A character count depends on the counting convention: JavaScript's string length counts UTF-16 code units, while a user-perceived grapheme can contain several code points. A UTF-8 byte count measures encoded storage. An ASCII letter uses one UTF-8 byte; many other characters use more. Neither count tells you how a model tokenizer segments the text.
A token is one unit produced by a particular tokenizer. It may represent part of a word, a whole word, punctuation, whitespace, or a byte-related fragment. Token count depends on the exact text and tokenizer. A context window is a token capacity for a particular model configuration; prompt tokens, instructions, retrieved material, tool messages, and generated output may all consume it.
Visible characters invite more than one counting convention.
Four ASCII bytes, two for é, one space, and four for the emoji.
Must be measured with the tokenizer used by the model.
Input and output compete for available capacity.
Declare the ratio and call the result an estimate.
For an early estimate, suppose ordinary English prose averages about four characters per
token. If a prompt has 8,000 counted characters, the estimate is ceil(8,000 / 4) = 2,000 tokens. The ceiling avoids understating a partial token in this simple arithmetic. The ratio is
a planning assumption, not a tokenizer guarantee.
For content that mixes source code, URLs, tables, numbers, multilingual text, or emoji, sample representative requests and measure them with the target tokenizer. Better still, collect actual usage after a pilot, because prompt templates and retrieved context often change while a product is being built.
| Input | Assumption | Calculation | Estimate |
|---|---|---|---|
| 8,000 plain-English characters | 4 characters/token | ceil(8,000 ÷ 4) | 2,000 tokens |
| 2,000 generated tokens | Requested output cap reached | Count separately | 2,000 tokens |
Multiply each token class by its own rate, then by the workload.
At the illustrative rates of $2 per million input tokens and $8 per million output tokens,
a request with 2,000 input and 2,000 output tokens costs (2,000 × $2 + 2,000 × $8) / 1,000,000 = $0.02. At 500 requests each day for 30 days, that scenario is $0.02 × 500 × 30 = $300. This arithmetic ignores discounts, minimums, caching
rules, retries, tools priced separately, and requests with different lengths.
Use the workload population that matches the decision: average requests for a rough monthly estimate, peak concurrent requests for capacity, and a high percentile or conservative envelope when a cost ceiling must not be exceeded. Avoid multiplying an average input by an unrelated worst-case output and calling it a measured monthly bill; label scenarios clearly.
$0.004 per request.
$0.016 per request.
Input and output charges combined.
500 requests/day × 30 days.
A request can fit its cost estimate and still exceed its token budget.
For an illustrative 16,000-token window, reserve 1,500 tokens for system instructions,
formatting, and tool overhead. If the estimated prompt takes 2,000 tokens and the response
could use 2,000, the remaining budget is 16,000 − 1,500 − 2,000 − 2,000 = 10,500 tokens. That remainder is available for additional prompt material only if the chosen model's
accounting and request structure match the assumptions.
Do not treat the entire remainder as safe input. Keep room for variable framing, retrieval, and output; check whether the model counts generated tokens against the same limit; and verify what happens on overflow. A context-limit error, truncation, or shorter response are different operational outcomes.
See how text, volume, and rate assumptions move the forecast.
Enter a representative prompt, then edit the rough characters-per-token assumption and editable unit rates. The lesson counts Unicode code points and UTF-8 bytes for comparison, but token estimation uses the declared character ratio. It does not call a real tokenizer.
Keep the token counts, unit rates, and request volume separate.
The examples calculate an input estimate from a stated character ratio, price input and output separately per million tokens, multiply by request volume, and calculate remaining context. They do not tokenize text or fetch model pricing. Go uses integer token counts and floating-point currency arithmetic; a production billing system should preserve the provider's reported precision and rounding rules.
The inputs are illustrative, and the rate values are placeholders.
export type Estimate = {
inputTokens: number;
outputTokens: number;
costPerRequest: number;
totalCost: number;
contextRemaining: number;
};
// Character-based estimate for a declared plain-English approximation; this is not a tokenizer.
export function estimate(
inputCharacters: number,
charsPerToken: number,
outputTokens: number,
inputUsdPerMillion: number,
outputUsdPerMillion: number,
requests: number,
contextWindow: number,
reservedTokens: number
): Estimate {
const values = [
inputCharacters,
charsPerToken,
outputTokens,
inputUsdPerMillion,
outputUsdPerMillion,
requests,
contextWindow,
reservedTokens
];
if (
!values.every(Number.isFinite) ||
inputCharacters < 0 ||
charsPerToken <= 0 ||
outputTokens < 0 ||
inputUsdPerMillion < 0 ||
outputUsdPerMillion < 0 ||
requests < 0 ||
contextWindow < 0 ||
reservedTokens < 0
) {
throw new RangeError(
'counts and rates must be finite and non-negative; charsPerToken must be positive'
);
}
if (
![inputCharacters, outputTokens, requests, contextWindow, reservedTokens].every(
Number.isInteger
)
) {
throw new RangeError(
'character and token counts, requests, and context sizes must be integers'
);
}
const inputTokens = Math.ceil(inputCharacters / charsPerToken);
const costPerRequest =
(inputTokens * inputUsdPerMillion) / 1_000_000 +
(outputTokens * outputUsdPerMillion) / 1_000_000;
return {
inputTokens,
outputTokens,
costPerRequest,
totalCost: costPerRequest * requests,
contextRemaining: contextWindow - reservedTokens - inputTokens - outputTokens
};
}
// Illustrative scenario: 8,000 prompt characters, 2,000 output tokens, 500 requests.
const run = estimate(8_000, 4, 2_000, 2, 8, 500, 16_000, 1_500);
console.log(run);
package main
import (
"errors"
"fmt"
"math"
)
type Estimate struct {
InputTokens int64
OutputTokens int64
CostPerRequestUSD float64
TotalCostUSD float64
ContextRemaining int64
}
// Character-based estimate for a declared plain-English approximation; this is not a tokenizer.
func EstimateCost(inputCharacters, charsPerToken, outputTokens, inputUSDPerMillion,
outputUSDPerMillion, requests, contextWindow, reservedTokens float64) (Estimate, error) {
values := []float64{inputCharacters, charsPerToken, outputTokens, inputUSDPerMillion,
outputUSDPerMillion, requests, contextWindow, reservedTokens}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
return Estimate{}, errors.New("counts and rates must be finite and non-negative")
}
}
for _, count := range []float64{inputCharacters, outputTokens, requests, contextWindow, reservedTokens} {
if math.Trunc(count) != count {
return Estimate{}, errors.New("character and token counts, requests, and context sizes must be integers")
}
}
if charsPerToken == 0 {
return Estimate{}, errors.New("charsPerToken must be positive")
}
input := int64(math.Ceil(inputCharacters / charsPerToken))
output := int64(outputTokens)
perRequest := float64(input)*inputUSDPerMillion/1_000_000 +
float64(output)*outputUSDPerMillion/1_000_000
return Estimate{
InputTokens: input,
OutputTokens: output,
CostPerRequestUSD: perRequest,
TotalCostUSD: perRequest * requests,
ContextRemaining: int64(contextWindow) - int64(reservedTokens) - input - output,
}, nil
}
func main() {
// Rates are placeholders entered by the operator, not a provider quote.
run, err := EstimateCost(8_000, 4, 2_000, 2, 8, 500, 16_000, 1_500)
if err != nil {
panic(err)
}
fmt.Printf("estimated input=%d output=%d cost/request=$%.5f total=$%.2f context remaining=%d\n",
run.InputTokens, run.OutputTokens, run.CostPerRequestUSD, run.TotalCostUSD, run.ContextRemaining)
}
A forecast helps plan a measurement; usage records support the next decision.
A rough ratio is useful when there is no tokenizer integration and only an early prompt draft. It becomes less useful when text includes unusual Unicode, source code, repeated templates, large retrieval results, or tool schemas. Measure representative serialized requests with the exact tokenizer/model path, then compare estimates with reported usage.
For a pilot, record request-level input and output token counts, model/version, task type, retries, tool usage, and failures. Compare medians and upper percentiles as well as total spend; a daily mean can hide occasional very long requests. Investigate why an outlier occurred before treating it as noise. A healthy forecast makes the next uncertainty visible instead of presenting a guessed number as a bill.
- OpenAI Tokenizer — an example of inspecting how a specific tokenizer segments text; different models can use different tokenization.
- OpenAI API latency optimization guide — operational context for reducing prompt and generated-token work; verify current API behavior and pricing separately.