← Math in Practice
Concept Math behind AI

Token and cost arithmetic

Budget a request with explicit assumptions, then replace the estimate with measured usage before making a commitment.

A team is considering an assistant for 500 incident summaries each day. The draft estimate multiplies prompt length by a per-token rate, but the actual tokenizer, system instructions, retrieved notes, tool calls, and response size are not in the spreadsheet yet. Before anyone promises a monthly budget, make the units and assumptions visible.

The judgment to keep

Characters and UTF-8 bytes describe text; tokens are the tokenizer's units; the context window is a model-specific token budget. Estimate input and output separately using editable unit rates, add the repeated-request volume, and leave headroom. A character ratio can help with an early forecast, but only the actual tokenizer or measured request usage can confirm token counts.

TypeScriptGo Characters · UTF-8 bytes · tokens · context window · input/output rates · repeated-request cost
01 / Read the budget forecast

A plausible monthly number can still rest on the wrong denominator.

The incident-response team expects 500 summaries a day. A spreadsheet assumes 2,000 input tokens and 500 output tokens per request, then applies one rate to the total. That forecast may be off because input and output can have different rates, and because the prompt may grow when incident notes or tool results are added. The arithmetic is only as sound as its counted tokens and rate units.

The example rates in this lesson are editable placeholders, not a quote from a vendor. We will assume $2 per million input tokens and $8 per million output tokens for the arithmetic. Replace them with the current rates for the exact model and billing path before using a real forecast.

Case file / Incident summary pilotEstimate 500 daily requests before the team chooses a monthly budget.
Known
Expected request count, a draft prompt, and a target response format.
Unknown
Exact tokenizer counts, variable retrieved context, retries, tool calls, and actual output lengths.
Decision
Is an order-of-magnitude forecast enough for a pilot, and what must be measured before scaling?
02 / Separate text units

A character, a byte, a token, and a context slot answer different questions.

A character count depends on the counting convention: JavaScript's string length counts UTF-16 code units, while a user-perceived grapheme can contain several code points. A UTF-8 byte count measures encoded storage. An ASCII letter uses one UTF-8 byte; many other characters use more. Neither count tells you how a model tokenizer segments the text.

A token is one unit produced by a particular tokenizer. It may represent part of a word, a whole word, punctuation, whitespace, or a byte-related fragment. Token count depends on the exact text and tokenizer. A context window is a token capacity for a particular model configuration; prompt tokens, instructions, retrieved material, tool messages, and generated output may all consume it.

Text“café 🧪”

Visible characters invite more than one counting convention.

UTF-811 bytes

Four ASCII bytes, two for é, one space, and four for the emoji.

TokensTokenizer-specific

Must be measured with the tokenizer used by the model.

ContextBudgeted tokens

Input and output compete for available capacity.

03 / Make an explicit estimate

Declare the ratio and call the result an estimate.

For an early estimate, suppose ordinary English prose averages about four characters per token. If a prompt has 8,000 counted characters, the estimate is ceil(8,000 / 4) = 2,000 tokens. The ceiling avoids understating a partial token in this simple arithmetic. The ratio is a planning assumption, not a tokenizer guarantee.

For content that mixes source code, URLs, tables, numbers, multilingual text, or emoji, sample representative requests and measure them with the target tokenizer. Better still, collect actual usage after a pilot, because prompt templates and retrieved context often change while a product is being built.

One rough estimate; assumed characters-per-token ratio is explicitly stated.
InputAssumptionCalculationEstimate
8,000 plain-English characters4 characters/tokenceil(8,000 ÷ 4)2,000 tokens
2,000 generated tokensRequested output cap reachedCount separately2,000 tokens
04 / Calculate repeated-request cost

Multiply each token class by its own rate, then by the workload.

At the illustrative rates of $2 per million input tokens and $8 per million output tokens, a request with 2,000 input and 2,000 output tokens costs (2,000 × $2 + 2,000 × $8) / 1,000,000 = $0.02. At 500 requests each day for 30 days, that scenario is $0.02 × 500 × 30 = $300. This arithmetic ignores discounts, minimums, caching rules, retries, tools priced separately, and requests with different lengths.

Use the workload population that matches the decision: average requests for a rough monthly estimate, peak concurrent requests for capacity, and a high percentile or conservative envelope when a cost ceiling must not be exceeded. Avoid multiplying an average input by an unrelated worst-case output and calling it a measured monthly bill; label scenarios clearly.

Input2,000 × $2/M

$0.004 per request.

Output2,000 × $8/M

$0.016 per request.

Per request$0.020

Input and output charges combined.

Monthly scenario$300

500 requests/day × 30 days.

05 / Check the context budget

A request can fit its cost estimate and still exceed its token budget.

For an illustrative 16,000-token window, reserve 1,500 tokens for system instructions, formatting, and tool overhead. If the estimated prompt takes 2,000 tokens and the response could use 2,000, the remaining budget is 16,000 − 1,500 − 2,000 − 2,000 = 10,500 tokens. That remainder is available for additional prompt material only if the chosen model's accounting and request structure match the assumptions.

Do not treat the entire remainder as safe input. Keep room for variable framing, retrieval, and output; check whether the model counts generated tokens against the same limit; and verify what happens on overflow. A context-limit error, truncation, or shorter response are different operational outcomes.

Sanity check: Prompt + reserved overhead + requested output must fit inside the configured window. If the result is negative, the scenario cannot fit under these assumptions even before a safety margin.
06 / Change the workload

See how text, volume, and rate assumptions move the forecast.

Enter a representative prompt, then edit the rough characters-per-token assumption and editable unit rates. The lesson counts Unicode code points and UTF-8 bytes for comparison, but token estimation uses the declared character ratio. It does not call a real tokenizer.

Code points180
UTF-8 bytes180
Estimated input tokens45
Input estimate45 tokens
Per-request cost$0.00409
Scenario total$61.35
Context remaining13,955 tokens
07 / Practice in code

Keep the token counts, unit rates, and request volume separate.

The examples calculate an input estimate from a stated character ratio, price input and output separately per million tokens, multiply by request volume, and calculate remaining context. They do not tokenize text or fetch model pricing. Go uses integer token counts and floating-point currency arithmetic; a production billing system should preserve the provider's reported precision and rounding rules.

Compare the same estimate in TypeScript and Go.

The inputs are illustrative, and the rate values are placeholders.

TypeScriptEstimate token cost and remaining context
token-cost.ts
export type Estimate = {
	inputTokens: number;
	outputTokens: number;
	costPerRequest: number;
	totalCost: number;
	contextRemaining: number;
};

// Character-based estimate for a declared plain-English approximation; this is not a tokenizer.
export function estimate(
	inputCharacters: number,
	charsPerToken: number,
	outputTokens: number,
	inputUsdPerMillion: number,
	outputUsdPerMillion: number,
	requests: number,
	contextWindow: number,
	reservedTokens: number
): Estimate {
	const values = [
		inputCharacters,
		charsPerToken,
		outputTokens,
		inputUsdPerMillion,
		outputUsdPerMillion,
		requests,
		contextWindow,
		reservedTokens
	];
	if (
		!values.every(Number.isFinite) ||
		inputCharacters < 0 ||
		charsPerToken <= 0 ||
		outputTokens < 0 ||
		inputUsdPerMillion < 0 ||
		outputUsdPerMillion < 0 ||
		requests < 0 ||
		contextWindow < 0 ||
		reservedTokens < 0
	) {
		throw new RangeError(
			'counts and rates must be finite and non-negative; charsPerToken must be positive'
		);
	}
	if (
		![inputCharacters, outputTokens, requests, contextWindow, reservedTokens].every(
			Number.isInteger
		)
	) {
		throw new RangeError(
			'character and token counts, requests, and context sizes must be integers'
		);
	}
	const inputTokens = Math.ceil(inputCharacters / charsPerToken);
	const costPerRequest =
		(inputTokens * inputUsdPerMillion) / 1_000_000 +
		(outputTokens * outputUsdPerMillion) / 1_000_000;
	return {
		inputTokens,
		outputTokens,
		costPerRequest,
		totalCost: costPerRequest * requests,
		contextRemaining: contextWindow - reservedTokens - inputTokens - outputTokens
	};
}

// Illustrative scenario: 8,000 prompt characters, 2,000 output tokens, 500 requests.
const run = estimate(8_000, 4, 2_000, 2, 8, 500, 16_000, 1_500);
console.log(run);
GoEstimate token cost and remaining context
token-cost.go
package main

import (
	"errors"
	"fmt"
	"math"
)

type Estimate struct {
	InputTokens       int64
	OutputTokens      int64
	CostPerRequestUSD float64
	TotalCostUSD      float64
	ContextRemaining  int64
}

// Character-based estimate for a declared plain-English approximation; this is not a tokenizer.
func EstimateCost(inputCharacters, charsPerToken, outputTokens, inputUSDPerMillion,
	outputUSDPerMillion, requests, contextWindow, reservedTokens float64) (Estimate, error) {
	values := []float64{inputCharacters, charsPerToken, outputTokens, inputUSDPerMillion,
		outputUSDPerMillion, requests, contextWindow, reservedTokens}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
			return Estimate{}, errors.New("counts and rates must be finite and non-negative")
		}
	}
	for _, count := range []float64{inputCharacters, outputTokens, requests, contextWindow, reservedTokens} {
		if math.Trunc(count) != count {
			return Estimate{}, errors.New("character and token counts, requests, and context sizes must be integers")
		}
	}
	if charsPerToken == 0 {
		return Estimate{}, errors.New("charsPerToken must be positive")
	}
	input := int64(math.Ceil(inputCharacters / charsPerToken))
	output := int64(outputTokens)
	perRequest := float64(input)*inputUSDPerMillion/1_000_000 +
		float64(output)*outputUSDPerMillion/1_000_000
	return Estimate{
		InputTokens:       input,
		OutputTokens:      output,
		CostPerRequestUSD: perRequest,
		TotalCostUSD:      perRequest * requests,
		ContextRemaining:  int64(contextWindow) - int64(reservedTokens) - input - output,
	}, nil
}

func main() {
	// Rates are placeholders entered by the operator, not a provider quote.
	run, err := EstimateCost(8_000, 4, 2_000, 2, 8, 500, 16_000, 1_500)
	if err != nil {
		panic(err)
	}
	fmt.Printf("estimated input=%d output=%d cost/request=$%.5f total=$%.2f context remaining=%d\n",
		run.InputTokens, run.OutputTokens, run.CostPerRequestUSD, run.TotalCostUSD, run.ContextRemaining)
}
08 / Replace estimates with evidence

A forecast helps plan a measurement; usage records support the next decision.

A rough ratio is useful when there is no tokenizer integration and only an early prompt draft. It becomes less useful when text includes unusual Unicode, source code, repeated templates, large retrieval results, or tool schemas. Measure representative serialized requests with the exact tokenizer/model path, then compare estimates with reported usage.

For a pilot, record request-level input and output token counts, model/version, task type, retries, tool usage, and failures. Compare medians and upper percentiles as well as total spend; a daily mean can hide occasional very long requests. Investigate why an outlier occurred before treating it as noise. A healthy forecast makes the next uncertainty visible instead of presenting a guessed number as a bill.

  • OpenAI Tokenizer — an example of inspecting how a specific tokenizer segments text; different models can use different tokenization.
  • OpenAI API latency optimization guide — operational context for reducing prompt and generated-token work; verify current API behavior and pricing separately.