← Math in Practice
Concept Text and representation

Text, Unicode, and encoded bytes

A visible character does not promise a fixed number of bytes.

A profile import rejects a display name that looks four characters long. The database field allows four bytes, and the name is café. The visible text is short, but its final letter takes more than one byte in UTF-8. Trace what the system measures before changing the limit or trimming the value.

The judgment to keep

Ask what unit the boundary measures: bytes, code points, UTF-16 code units, or user-perceived characters. Encode or segment with the matching rule, and preserve complete characters at the boundary.

TypeScriptGo Unicode code points · UTF-8 · byte limits · grapheme clusters
01 / Read the truncation report

The database limit is measured in bytes, while the UI shows characters.

A profile service accepts display names from browsers and stores them in a legacy field with a four-byte UTF-8 limit. The name café has four Unicode code points, but its UTF-8 encoding is five bytes: c, a, and f each use one byte; é uses two. A function that slices the first four bytes can split that last encoded character.

This example gives the storage boundary a clear contract: accept only a valid UTF-8 prefix that fits within four bytes. That does not automatically define a good product rule for names. The limit may be a legacy constraint, and a user-facing character limit should be defined separately from storage. The sample data here is illustrative; measure the actual field and the exact serialization path before deciding what failed.

Case file / Profile importDetermine whether the name was rejected, corrupted, or clipped.
Input
café, four Unicode code points.
Storage contract
At most 4 UTF-8 bytes, valid encoding required.
Observed clue
ASCII names pass; some accented names fail or appear incomplete.
Question
Which operation can split the encoding, and what should the diagnostic log record?
Keep: the field's encoding, the unit of its limit, and whether the operation rejects or truncates.
02 / Separate the text units

“Length” can answer several different questions.

A byte is an 8-bit storage unit. A Unicode code point is a number assigned to a character or control in the Unicode code space. An encoding such as UTF-8 maps code points to bytes. A grapheme cluster is an approximation of one user-perceived character; it may contain multiple code points, such as a base letter followed by a combining accent or an emoji joined with a modifier.

For the simple text café, there are 4 code points and 5 UTF-8 bytes. In JavaScript, text.length counts UTF-16 code units, which can differ again: an emoji outside the Basic Multilingual Plane takes two UTF-16 code units but one code point. In Go, len(text) counts bytes; utf8.RuneCountInString(text) counts decoded code points. Neither code-point count is a general grapheme-cluster count.

So write the question before reaching for a length function. A protocol frame asks for bytes. A code-point iteration asks for Unicode scalar values. A UI character counter may need grapheme segmentation. The same string can correctly have different lengths under those contracts.

Illustrative counts · UTF-8 encoding · grapheme examples depend on Unicode segmentation rules
TextUTF-8 bytesCode pointsGrapheme clustersWhy counts differ
cafe444Four ASCII code points, one byte each.
café544The precomposed é uses two UTF-8 bytes.
é321e plus a combining acute accent.
🙂411One code point encoded as four UTF-8 bytes.
03 / Follow UTF-8 encoding

UTF-8 uses one to four bytes for a Unicode code point.

UTF-8 represents code points with byte sequences of different lengths. ASCII values in the range U+0000–U+007F retain their familiar single-byte representation. Other values use multiple bytes, up to four. This variable length makes ordinary English-heavy text compact while allowing the same encoding to represent the full Unicode range.

The number of bytes is not a visual-width measure. A four-byte emoji may render narrow or wide depending on font and platform. A grapheme cluster may contain several code points, and a text editor can display that sequence as one user-perceived unit. Storage, cursor movement, deletion, display width, and user-facing “characters” are different operations and can need different APIs.

For café, UTF-8 bytes can be inspected as hexadecimal: 63 61 66 C3 A9. The limit of four ends between C3 and A9. A decoder cannot interpret that incomplete prefix as the original é. A safe storage boundary either rejects the whole value or truncates at a valid encoding boundary according to an explicit product rule.

01 / TEXTcafé

Four code points as read by a person.

02 / ENCODE63 61 66 C3 A9

Five bytes in UTF-8.

03 / LIMIT4 bytes

Boundary lands inside the two-byte encoding of é.

04 / SAFE PREFIXcaf

Three bytes; the next complete code point needs two.

04 / Diagnose the boundary

Locate the transformation where the first bad byte appears.

When text arrives corrupted, do not assume the database caused it. Trace the value at each boundary: request decoding, application validation, normalization, serialization, database driver conversion, column storage, query decoding, and UI rendering. Compare safe measurements at each stage: declared encoding, byte length, validity, and whether the text changed. A value can be valid when received and become invalid only after a byte slice, or remain intact in storage while a UI counter misreports its length.

Use paired probes that differ in one property: cafe versus café tests multibyte encoding; precomposed é versus e plus combining accent tests normalization; 🙂 tests both UTF-8 byte length and JavaScript UTF-16 code-unit behavior. Keep these test fixtures explicit so a future refactor does not silently change the assumption under test.

Ask what the field is supposed to store. If it stores display text, a user-facing limit should usually count grapheme clusters and storage should still allow enough encoded bytes. If it stores protocol bytes, character-aware normalization may alter the payload and is likely inappropriate. If it stores an identifier, the accepted normalization and comparison rules belong in its specification.

05 / Practice in code

Measure UTF-8 bytes and preserve complete code points.

The TypeScript example uses the browser's TextEncoder to count UTF-8 bytes, then iterates by Unicode code point while keeping a prefix under the byte budget. The Go example validates the string as UTF-8, uses len for bytes and utf8.RuneCountInString for code points, then chooses a rune-aligned prefix.

Both implementations are deliberate but incomplete for a user-facing character limit: neither guarantees the result ends at a grapheme-cluster boundary. TypeScript strings can contain unpaired UTF-16 surrogates; this example rejects them before encoding, since TextEncoder would otherwise encode the replacement character. The Go example rejects invalid UTF-8 before slicing. A production contract should decide how malformed input is handled.

Compare the same byte-boundary problem in TypeScript and Go.

Both examples count UTF-8 bytes and return only complete encoded code points within the budget.

TypeScriptText and bytes · UTF-8 length and safe prefix
text.ts
const encoder = new TextEncoder();

export function utf8ByteLength(text: string): number {
	return encoder.encode(text).byteLength;
}

/** Keep complete Unicode code points under a UTF-8 byte budget.
 * This does not preserve whole grapheme clusters such as emoji plus modifiers. */
export function prefixWithinUtf8Bytes(text: string, maxBytes: number): string {
	if (!Number.isSafeInteger(maxBytes) || maxBytes < 0) {
		throw new Error('maxBytes must be a non-negative safe integer');
	}
	let prefix = '';
	let usedBytes = 0;
	for (const codePoint of text) {
		const firstUnit = codePoint.charCodeAt(0);
		if (codePoint.length === 1 && firstUnit >= 0xd800 && firstUnit <= 0xdfff) {
			throw new Error('text contains an unpaired UTF-16 surrogate');
		}
		const codePointBytes = utf8ByteLength(codePoint);
		if (usedBytes + codePointBytes > maxBytes) break;
		prefix += codePoint;
		usedBytes += codePointBytes;
	}
	return prefix;
}

const name = 'café';
console.log({ text: name, bytes: utf8ByteLength(name), codePoints: [...name].length });
// { text: 'café', bytes: 5, codePoints: 4 }
console.log(prefixWithinUtf8Bytes(name, 4));
// caf — the next complete code point, é, needs two UTF-8 bytes.
GoText and bytes · UTF-8 length and safe prefix
text.go
package main

import (
	"fmt"
	"unicode/utf8"
)

// PrefixByBytes keeps complete UTF-8 encoded runes within a byte budget.
// It counts Unicode code points, not user-perceived grapheme clusters.
func PrefixByBytes(text string, maxBytes int) (string, error) {
	if maxBytes < 0 {
		return "", fmt.Errorf("maxBytes must be non-negative")
	}
	if !utf8.ValidString(text) {
		return "", fmt.Errorf("input is not valid UTF-8")
	}
	used := 0
	for _, r := range text {
		runeBytes := utf8.RuneLen(r)
		if used+runeBytes > maxBytes {
			break
		}
		used += runeBytes
	}
	return text[:used], nil
}

func main() {
	name := "café"
	fmt.Printf("text=%q bytes=%d code-points=%d\n", name, len(name), utf8.RuneCountInString(name))
	// text="café" bytes=5 code-points=4

	prefix, err := PrefixByBytes("café", 4)
	if err != nil {
		panic(err)
	}
	fmt.Printf("4-byte prefix=%q bytes=%d\n", prefix, len(prefix))
	// 4-byte prefix="caf" bytes=3; the next é needs two bytes.
}
06 / Change the constraint

Pick the unit that matches the user promise and the system boundary.

Practice / Display-name policyThe product wants a maximum of 12 visible characters and a database column can hold 48 UTF-8 bytes.
Task A
Explain why twelve code points and twelve grapheme clusters may not be equivalent.
Task B
Decide how to enforce the 48-byte storage boundary without creating invalid UTF-8.
Transfer
What changes if the same field is a signed protocol token rather than display text?
Show a worked answer

A product promise about visible characters should normally be implemented using a Unicode grapheme segmentation algorithm, because one displayed unit can contain several code points. A storage limit is still bytes after encoding. Validate or segment the text according to the product rule, encode it as UTF-8, measure the encoded bytes, and reject or safely truncate if it exceeds 48 bytes. Do not silently assume each grapheme costs one byte.

For an opaque signed protocol token, preserve the exact byte sequence covered by the signature. Do not normalize or truncate it as display text. Validate against the protocol's byte limit and reject an oversized token unless the protocol defines a different operation.

Take with you: text is content; UTF-8 bytes are one representation of it. Name the unit before counting, and never split a multibyte sequence by accident.

Further reading: Unicode FAQ on UTF encodings, Unicode FAQ on characters and combining marks, and the Go unicode/utf8 package documentation. See also the ECMAScript specification for String length and WHATWG's TextEncoder interface.