← Math in Practice
Concept Quantities and representation

Entropy and bits of strength

Length gives a useful estimate only when you know how each choice was made.

A product team wants to email an invitation link. The token must be easy to transport and unrealistic to guess, but the review has only one proposed number: “It is 20 characters long.” That length cannot answer the security question by itself. We need to know the alphabet, how each character is selected, and what an attacker can try or observe.

The judgment to keep

Treat entropy as a model of a probability distribution. For independently and uniformly selected characters, calculate length × log₂(alphabet size). Then check whether the generator, token handling, and attack conditions support that model.

TypeScriptGo Uniform outcome spaces · entropy in bits · token design · generation assumptions · threat models
01 / Choose a token

The important question is how hard it is to land on a valid secret.

A service issues an invitation token to grant access to one workspace. It is valid for 24 hours, can be redeemed once, and is accepted by an endpoint that limits attempts. The team proposes a 20-character token drawn from a 32-symbol alphabet. We will use this as an explicit, illustrative model: each of the 32 symbols is chosen with equal probability at each position, independently of the other positions.

The alphabet might be a carefully selected set of letters and digits that avoids confusing characters in a copied URL. But the human-readable design choice is only one part of the model. A secure random source and unbiased selection must make every allowed choice equally likely. A value made with a timestamp, a counter, a user name, or a general-purpose pseudorandom generator does not inherit the estimate just because it has the same visible length.

Case file / Workspace invitationWhat does “20 characters” actually buy?
Purpose
Bearer token grants one-time workspace access.
Lifetime
24 hours, then the invitation expires.
Proposed model
20 positions; 32 equally likely symbols per position.
Decision
Check the outcome count and the assumptions behind it.
02 / Count the outcomes

Bits turn a number of equally likely possibilities into a base-two scale.

One fair binary choice has two possible outcomes and represents one bit. A uniform choice among N outcomes represents log₂(N) bits because 2log₂(N) = N. Bits are convenient because doubling the number of possible secrets adds one bit, regardless of how large the space already is.

For a length-L string, if each position independently chooses one of A allowed symbols with equal probability, multiplication gives N = AL. Taking base-two logs turns that product into a sum: H = log₂(AL) = L × log₂(A). This is the Shannon entropy of the stated uniform model. In security conversation it is often used as a rough guess-space measure; it is not a complete security score.

01 / Choices per positionA = 32

Every position has 32 equally likely symbols.

02 / PositionsL = 20

Choices are independent across positions.

03 / Model result20 × log₂(32) = 100 bits

Since 32 = 2⁵, each position contributes exactly 5 modeled bits.

The same arithmetic can be checked through the full search space: 32²⁰ = (2⁵)²⁰ = 2¹⁰⁰ possible strings. An attacker who can test q distinct candidates against a uniformly generated token succeeds with probability q / 2¹⁰⁰, provided q ≤ 2¹⁰⁰ and the candidate order gives no information. That condition is a mathematical model, not an estimate of real attack throughput or an endorsement of a particular bit threshold.

03 / Answer the review question

A bit count answers one narrow question about a secret's possible values.

04 / Inspect how it is generated

Entropy belongs to the process that produces values, not their printed shape.

The formula assumes every allowed string can occur and that each is equally likely. If the generator samples each position independently and uniformly, all AL strings meet that assumption. If it picks an index by taking a random byte modulo 32, the byte range happens to divide evenly by 32; for other alphabet sizes, modulo reduction can give some symbols more possible source bytes than others. A standard bounded-random API can handle this without a hand-built rejection sampler.

Dependence also changes the count. If the first position is selected from 10 templates and the remaining positions are derived from that template, the nominal character alphabet doesn't describe the reachable distribution. Predictable seeds and leaked state can narrow the possible outputs further. Conversely, cryptographic generators are designed to make outputs unpredictable to an attacker; the language's cryptographic API and the way it is used are part of the design evidence.

Practical distinction

A secret can be hard to guess and still be easy to steal.

Entropy estimates the uncertainty of a value before an attacker learns anything about it. It does not protect a token copied into analytics, exposed in a referrer header, reused across accounts, or accepted forever. For passwords, an offline attacker who steals hashes faces a different problem from an online attacker constrained by login throttling; storage and rate limits change the threat model, not the original generation distribution.

05 / Change the model

Explore the estimate, while keeping the uniformity assumption in view.

Set a number of symbols per position and the string length. The calculator reports only the modeled entropy for independent uniform choices; it does not inspect a real password, sample a generator, or decide whether a token is safe to deploy.

Model result 100.00 bits

20 × log₂(32) = 100.00

Assumption: every symbol is equally likely and each position is independent.

Try 16 characters from 64 equally likely choices: 16 × log₂(64) = 16 × 6 = 96 bits. Now imagine the generator chooses one of 10 templates, then fills only three positions uniformly from those 64 symbols. The original 16 × 6 calculation no longer follows: repeated positions are constrained by a shared template. To estimate the true distribution, you need the generation procedure, not just the displayed alphabet and length.

06 / Practice in code

Make the assumptions visible in a small calculation.

These TypeScript and Go examples calculate the same modeled quantity and reject invalid inputs. They do not generate secrets: use the runtime's cryptographic random facilities for production generation, and use an unbiased bounded-integer operation when selecting from an alphabet. Node's crypto.randomInt documents an unbiased range; Go's crypto/rand package provides cryptographically secure random values and a uniform bounded-integer function.

Compare the same model in TypeScript and Go.

Both examples calculate length × log₂(alphabet size); neither one generates a secret.

TypeScriptUniform-model entropy calculation
entropy.ts
/**
 * Entropy estimate for a fixed-length string whose characters are selected
 * independently and uniformly from the stated alphabet.
 * This models the generation process; it does not estimate human choices.
 */
export function uniformStringBits(alphabetSize: number, length: number): number {
	if (!Number.isSafeInteger(alphabetSize) || alphabetSize < 2) {
		throw new RangeError('alphabetSize must be a safe integer of at least 2');
	}
	if (!Number.isSafeInteger(length) || length < 1) {
		throw new RangeError('length must be a positive safe integer');
	}

	return length * Math.log2(alphabetSize);
}

const alphabetSize = 32;
const length = 20;
const bits = uniformStringBits(alphabetSize, length);

console.log(`${length} symbols from ${alphabetSize} choices each: ${bits} bits`);
GoUniform-model entropy calculation
entropy.go
package main

import (
	"fmt"
	"math"
)

// uniformStringBits estimates a fixed-length string's entropy when each
// character is selected independently and uniformly from the stated alphabet.
// It models the generation process; it does not estimate human choices.
func uniformStringBits(alphabetSize, length int) (float64, error) {
	if alphabetSize < 2 {
		return 0, fmt.Errorf("alphabetSize must be at least 2")
	}
	if length < 1 {
		return 0, fmt.Errorf("length must be positive")
	}
	return float64(length) * math.Log2(float64(alphabetSize)), nil
}

func main() {
	const alphabetSize = 32
	const length = 20

	bits, err := uniformStringBits(alphabetSize, length)
	if err != nil {
		panic(err)
	}
	fmt.Printf("%d symbols from %d choices each: %.0f bits\n", length, alphabetSize, bits)
}
07 / Diagnose the threat

Start with what the attacker can do, then decide which calculation applies.

  1. Name the value. Is it a public identifier, a password chosen by a person, or a bearer secret that grants access?
  2. Recover the generation process. Record the source of randomness, alphabet, length, selection method, seeding, and any structural constraints.
  3. Check the model. Only apply L × log₂(A) when allowed choices are independent and uniform at each position.
  4. State the attack boundary. Distinguish online attempts with throttling from offline guesses, and consider token theft, replay, and reuse separately.
  5. Choose evidence and controls. Inspect implementation and logs; then review expiry, one-time use, leakage paths, and consequences of compromise.
Transfer exercise

A person proposes a memorable 20-character passphrase.

Why is 20 × log₂(32) not a valid estimate merely because the password accepts 32 symbols? Describe the generation assumption that failed, then name two pieces of evidence or controls you would examine for online guessing and for a stolen password database. Separate your modeled calculation from what the actual password choice tells you.