← Concepts & practices
Technique Testing, debugging, and measurement

CPU and memory
profiling

Follow the evidence.

You can stare at a slow function for a long time and still fix the wrong thing. Let’s give the program a repeatable job, find where the work goes, and check what stays alive afterward.

The skill to keep

Turn “this feels slow or heavy” into an observation, a cause you can test, and a change you can defend.

TypeScriptGo A practical investigation
01 / Reproduce

Make “slow” specific enough to investigate.

A lesson catalog filters as someone searches. It works on a small dataset. With more records and a longer session, searches do more work and old result lists accumulate. We have two symptoms to investigate; we don’t yet have one explanation for both.

Profiling attributes resource use to parts of a running program. A timer tells us how long a replay took. A profile helps us investigate what consumed CPU or allocated memory during that work.

Case fileThe catalog gets heavier as you use it.
Job
Search a fixed catalog with the same query sequence.
Required behavior
Return every matching ID in input order, including all records for an empty query.
Starting implementation
Normalize each document on every search. Keep each result list in a debug history.
Product constraint
The view needs current results. Full search history is incidental debugging state, not an Undo feature.

Start with 5,000 generated records and 24 searches. Use the same dataset, query order, build, and environment for both versions. Separate setup from the repeated searches; building an index is work too.

Write down the symptom you want to improve before recording. For a real product, that might be typing-to-result latency or memory still held after closing a view. This lesson isolates search computation and result ownership; it does not measure rendering, requests, or a complete user interaction.

Leave this step with

A workload someone else can repeat, the behavior it must preserve, and the resource you are investigating.

02 / Capture

Ask one question of the right evidence.

Work

Where is CPU being spent?

Sample active call stacks during the repeated searches. If most elapsed time is waiting, follow that wait with a trace instead.

Creation

What keeps getting allocated?

Inspect allocation sites. Many temporary objects can create collection work without remaining alive.

Lifetime

What is still being kept?

Compare equivalent lifecycle points and find the owner of surviving objects. A large heap alone does not establish a leak.

Collect CPU and memory evidence separately so the instruments interfere less with one another. A profile is an observation under a workload, not a verdict on a function. Go’s diagnostics guide explains the distinction and the interference between tools.

Read the investigation in your languages.

Choose the source comparisons and capture recipes throughout this lesson.

Start here · browser

Capture the experiment on this page.

  1. Open DevTools → Performance. Start recording, run the comparison in step 4, then stop.
  2. Select a replay’s JavaScript task. Inspect the Main track, Bottom-up, and callers; exclude page load and the gaps between runs.
  3. Record which activity owns the work. In a production build, source maps help you connect generated code to its source.

The experiment includes output checks and warmup before its measured rounds. Those also appear in a recording. Separate setup, replay, and the surrounding interface. See Chrome’s Performance reference.

Capture retained memory in the browserA separate investigation
  1. Reset the lab and take a heap snapshot in DevTools → Memory.
  2. Choose “Memory: keep only current results,” run the comparison, then take another snapshot.
  3. Inspect the two live SearchSession instances and their history arrays. Both final sessions are held intentionally for inspection. Follow a result array’s retainers back to its session.
  4. Reset, take another snapshot, and check whether the session-owned objects disappear. Repeat at a larger search count to compare equivalent runs, rather than accumulating unrelated page activity.

The baseline keeps one list per search; the candidate keeps the last. This is a controlled retention example, not proof that the page itself leaks. Chrome snapshots trigger collection; console inspection can itself keep objects alive. A snapshot’s retaining path tells you who holds an object. See heap snapshots and retainers.

TypeScript outside the browserNode · CPU and allocation sampling

From this lesson’s examples/ directory, run the driver with Node 22.18 or later. The source uses supported erasable TypeScript syntax. Capture one resource at a time; open the CPU artifact in a compatible profile viewer and the heap artifact in DevTools Memory.

node-capture.sh
# From examples/ · Node 22.18+; run one profiler at a time
node --experimental-strip-types --cpu-prof --cpu-prof-name=baseline.cpuprofile profile.ts baseline
node --experimental-strip-types --cpu-prof --cpu-prof-name=index.cpuprofile profile.ts index

# Allocation sampling, in a separate process
node --experimental-strip-types --heap-prof --heap-prof-name=baseline.heapprofile profile.ts baseline

The two-second driver repeats fresh sessions. CPU capture includes startup and warmup: locate the search work. Allocation samples are not a heap snapshot or a retaining-path graph. Compare per-replay work as well as totals; faster versions execute more replays in the same interval. See the Node CPU flags and heap flags.

Copy the Node capture driverprofile.ts · alongside search.ts
profile.ts
// Node 22.18+; run separately with --cpu-prof OR --heap-prof.
import { performance } from 'node:perf_hooks';
import { makeDocuments, makeQueries, modes, replay, SearchSession, type Mode } from './search.ts';

const mode = process.argv[2] ?? 'baseline';
if (!modes.includes(mode as Mode)) throw new Error('Mode: baseline | index | release | both');
const documents = makeDocuments(10000);
const queries = makeQueries(48);
replay(new SearchSession(documents, mode as Mode), queries); // Warmup, still visible to startup profilers.
let held: SearchSession | undefined;
let iterations = 0;
const start = performance.now();
while (performance.now() - start < 2000) {
	held = new SearchSession(documents, mode as Mode);
	replay(held, queries);
	iterations++;
}
console.log({
	mode,
	iterations,
	elapsedMs: performance.now() - start,
	finalSession: held?.inspect()
});
// Fresh session per replay: this driver measures steady repeated work, not one endlessly growing history.
// Use the browser snapshot procedure to investigate retaining paths; a heap profile is not a snapshot.
Go: capture, then choose what pprof countsCPU · alloc_space · inuse_space

From examples/go/, use the included benchmark harness. Each operation is a fresh session plus 48 searches over 10,000 records. The final session stays referenced so its lifetime is defined in the heap capture.

go-capture.sh
# From examples/go/ · standard library only
go test -run '^$' -bench '^BenchmarkSearch/baseline$' -benchtime=2s -cpuprofile=baseline.cpu
go tool pprof -top baseline.cpu

# Separate capture: all allocation traffic versus still-live allocations
go test -run '^$' -bench '^BenchmarkSearch/baseline$' -benchtime=2s -memprofile=baseline.heap
go tool pprof -top -alloc_space baseline.heap
go tool pprof -top -inuse_space baseline.heap

# Repeat with /index, /release, or /both and a distinct output filename.

alloc_space counts sampled allocation volume, including objects already collected. inuse_space selects sampled live bytes. Both attribute allocations to sites; neither is a JavaScript-style retaining-path graph. CPU “flat” is work in the function; “cum” includes callees. Keep the sample type and workload in your note. See runtime/pprof.

Copy the Go tests and benchmark harnesssearch_test.go · alongside search.go and go.mod
search_test.go
package main

import (
	"encoding/json"
	"os"
	"reflect"
	"testing"
)

func TestCases(t *testing.T) {
	data, err := os.ReadFile("../cases.json")
	if err != nil {
		t.Fatal(err)
	}
	var cases []struct {
		Name      string
		Documents []Document
		Queries   []string
		Expected  [][]int
	}
	if err := json.Unmarshal(data, &cases); err != nil {
		t.Fatal(err)
	}
	for _, c := range cases {
		t.Run(c.Name, func(t *testing.T) {
			for _, mode := range []string{"baseline", "index", "release", "both"} {
				s := NewSession(c.Documents, mode)
				total := 0
				for i, q := range c.Queries {
					got := s.Search(q)
					if !reflect.DeepEqual(got, c.Expected[i]) {
						t.Fatalf("%s: %v != %v", mode, got, c.Expected[i])
					}
					total += len(got)
				}
				indexed := mode == "index" || mode == "both"
				keep := mode == "baseline" || mode == "index"
				norm, entries := len(c.Documents)*len(c.Queries), 0
				if indexed {
					norm, entries = len(c.Documents), len(c.Documents)
				}
				retained, batches := total, len(c.Queries)
				if !keep && batches > 0 {
					retained, batches = len(c.Expected[len(c.Expected)-1]), 1
				}
				want := Stats{norm, len(c.Documents) * len(c.Queries), total, retained, batches, entries}
				if s.Inspect() != want {
					t.Fatalf("%s: %+v != %+v", mode, s.Inspect(), want)
				}
			}
		})
	}
}

func TestSnapshotAndResults(t *testing.T) {
	for _, mode := range []string{"baseline", "index", "release", "both"} {
		docs := []Document{{4, "Factory"}}
		s := NewSession(docs, mode)
		docs[0].Text = "Stack"
		first := s.Search("factory")
		s.Search("stack")
		if !reflect.DeepEqual(first, []int{4}) {
			t.Fatal(first)
		}
	}
}

var held *SearchSession // Retain only the final session so inuse_space has a defined lifetime.
func BenchmarkSearch(b *testing.B) {
	for _, mode := range []string{"baseline", "index", "release", "both"} {
		b.Run(mode, func(b *testing.B) {
			docs, queries := MakeDocuments(10000), MakeQueries(48)
			b.ReportAllocs()
			b.ResetTimer()
			for i := 0; i < b.N; i++ {
				held = NewSession(docs, mode)
				if Replay(held, queries) != 195000 {
					b.Fatal("changed result")
				}
			}
		})
	}
}

Behavior tests read the shared ../cases.json. The profiling commands skip those tests with -run '^$'.

Leave this step with

A recording tied to one workload and a named metric: active CPU, allocation volume, or live memory.

03 / Interpret

A big frame is a place to look. Follow its callers.

A function can be prominent because each call is costly, because it is called often, or because its children do the work. Before rewriting it, find which of those explanations fits.

Reading practice · illustrative data100 CPU samples under a search replayAuthored to explain attribution. These are not measurements from your browser.
search replay100 inclusive · 10 self
normalize document60 self samples
match text25 self samples
record results5 self samples
replay’s own work10 self samples

The parent includes its descendants. Adding 100 + 60 + 25 + 5 would count the same samples twice. In a CPU flame graph, width represents accumulated samples; a timeline flame chart also preserves when events happened. Neither width nor color means “bad code.”

Our candidate hypothesis is specific: document text is normalized repeatedly even though the catalog has not changed. A real profile may group or inline that work differently. Connect the hot path to the repeated search before acting on it.

Memory needs another question. Result arrays are created by search, but retained by the session’s debug history. Changing the allocation site is one possibility; changing the owner’s lifetime is another. A counter that says “entries created” cannot tell you how many are still reachable.

Leave this step with

A suspected cause, the evidence connecting it to the symptom, and the result you expect from one change.

04 / Change

Make the smallest change your evidence supports.

Try precomputing first. Then run a separate comparison that changes only the history. Finally, combine them. Keeping those experiments apart lets us explain which decision caused which result.

Controlled experiment

Change one thing. Keep the search the same.

The TypeScript source runs here. Your language preferences change the code you read, not this runtime.

Predict: what work moves into setup, and what new state stays alive?

Choose one change, predict its effect, then run the comparison.

Evidence appears after a run.

Expect two views of the same workload: elapsed time, and a map of what the session still holds.

Counts describe this implementation’s operations and logical entries, not allocated bytes or a heap measurement. Timings include instrumentation, exclude rendering and data generation, and may overlap or reverse on small workloads. Both final sessions are held for inspection until reset, another run, or navigation.

Follow the change in the sourceOne workload · four modes
TypeScriptSearch and ownership
search.ts
export class SearchSession {
	private documents: Document[];
	private index: string[] | null;
	readonly history: number[][] = [];
	normalizedDocuments = 0;
	scannedDocuments = 0;
	createdResultSlots = 0;
	private keepHistory: boolean;

	constructor(documents: Document[], mode: Mode) {
		// Take a snapshot: edits require a new session and index.
		this.documents = documents.map((document) => ({ ...document }));
		this.keepHistory = mode === 'baseline' || mode === 'index';
		this.index =
			mode === 'index' || mode === 'both'
				? this.documents.map((document) => {
						this.normalizedDocuments++;
						return normalize(document.text);
					})
				: null;
	}

	search(query: string): number[] {
		const needle = normalize(query);
		const ids: number[] = [];
		for (let i = 0; i < this.documents.length; i++) {
			this.scannedDocuments++;
			let text = this.index?.[i];
			if (text === undefined) {
				this.normalizedDocuments++;
				text = normalize(this.documents[i].text);
			}
			if (text.includes(needle)) ids.push(this.documents[i].id);
		}
		this.createdResultSlots += ids.length;
		if (!this.keepHistory) this.history.length = 0;
		this.history.push(ids);
		return ids;
	}

	inspect() {
		return {
			normalizedDocuments: this.normalizedDocuments,
			scannedDocuments: this.scannedDocuments,
			createdResultSlots: this.createdResultSlots,
			retainedResultSlots: this.history.reduce((sum, ids) => sum + ids.length, 0),
			retainedBatches: this.history.length,
			indexEntries: this.index?.length ?? 0
		};
	}
}
GoSearch and ownership
search.go
type SearchSession struct {
	documents                                                 []Document
	index                                                     []string
	history                                                   [][]int
	keepHistory                                               bool
	normalizedDocuments, scannedDocuments, createdResultSlots int
}

func NewSession(documents []Document, mode string) *SearchSession {
	s := &SearchSession{documents: append([]Document(nil), documents...), keepHistory: mode == "baseline" || mode == "index"}
	if mode == "index" || mode == "both" {
		s.index = make([]string, len(documents))
		for i, document := range s.documents {
			s.index[i] = normalize(document.Text)
			s.normalizedDocuments++
		}
	}
	return s
}

func (s *SearchSession) Search(query string) []int {
	needle := normalize(query)
	ids := make([]int, 0)
	for i, document := range s.documents {
		s.scannedDocuments++
		var text string
		if s.index != nil {
			text = s.index[i]
		} else {
			text = normalize(document.Text)
			s.normalizedDocuments++
		}
		if strings.Contains(text, needle) {
			ids = append(ids, document.ID)
		}
	}
	s.createdResultSlots += len(ids)
	if !s.keepHistory {
		s.history = nil
	}
	s.history = append(s.history, ids)
	return ids
}

baseline repeats normalization and keeps every result list. index changes preparation only. release changes history only. both combines them. All modes snapshot the input catalog and preserve the same search results.

The counter records document normalizations, including index construction, but not query normalization. The search still visits every document: this is prepared text, not a lookup index that avoids scanning.

TypeScriptReplay the same queries
search.ts
export function replay(session: SearchSession, queries: string[]): number {
	let matches = 0;
	for (const query of queries) matches += session.search(query).length;
	return matches;
}

export function example(mode: Mode = 'baseline') {
	const session = new SearchSession(makeDocuments(1000), mode);
	const matches = replay(session, makeQueries(24));
	return { mode, matches, ...session.inspect() };
}
GoReplay the same queries
search.go
func Replay(s *SearchSession, queries []string) int {
	matches := 0
	for _, query := range queries {
		matches += len(s.Search(query))
	}
	return matches
}
func main() {
	for _, mode := range []string{"baseline", "index", "release", "both"} {
		s := NewSession(MakeDocuments(1000), mode)
		fmt.Println(mode, "matches:", Replay(s, MakeQueries(24)), "stats:", s.Inspect())
	}
}

The memory change deliberately gives up the debug archive. That is acceptable only because the stated product contract needs current results. If this were an audit log or an Undo history, deleting it would change the feature. We would need a different retention policy.

The CPU change also has a price: setup work, stored strings, and a freshness boundary. This sample snapshots a catalog. A live editable catalog needs an explicit rebuild or update policy; silently searching a stale index is not an optimization.

Language details that affect this investigation

TypeScript: the session holds JavaScript arrays; results are returned by reference and callers must treat them as read-only. Removing a session reference makes objects eligible for collection only if no other owner retains them.

Go: results use slices. Discarding the old history removes its references here; reslicing a larger buffer elsewhere can keep its backing allocation alive. The sample’s copied document structs still refer to immutable string data.

These differences are why logical entry counts are not cross-language byte measurements.

Complete source filesCopyable · no package dependencies
TypeScriptComplete search workload
search.ts
export type Document = { id: number; text: string };
export type Mode = 'baseline' | 'index' | 'release' | 'both';
export const modes: Mode[] = ['baseline', 'index', 'release', 'both'];
export const queryCycle = [
	'FACTORY',
	' service ',
	'stack',
	'missing',
	'',
	'go',
	'memory',
	'factory'
];

// ASCII-only fixture: no locale, Unicode case-folding, or search ranking contract.
export function normalize(text: string): string {
	return text.replace(/^[ \t\r\n]+|[ \t\r\n]+$/g, '').toLowerCase();
}

export class SearchSession {
	private documents: Document[];
	private index: string[] | null;
	readonly history: number[][] = [];
	normalizedDocuments = 0;
	scannedDocuments = 0;
	createdResultSlots = 0;
	private keepHistory: boolean;

	constructor(documents: Document[], mode: Mode) {
		// Take a snapshot: edits require a new session and index.
		this.documents = documents.map((document) => ({ ...document }));
		this.keepHistory = mode === 'baseline' || mode === 'index';
		this.index =
			mode === 'index' || mode === 'both'
				? this.documents.map((document) => {
						this.normalizedDocuments++;
						return normalize(document.text);
					})
				: null;
	}

	search(query: string): number[] {
		const needle = normalize(query);
		const ids: number[] = [];
		for (let i = 0; i < this.documents.length; i++) {
			this.scannedDocuments++;
			let text = this.index?.[i];
			if (text === undefined) {
				this.normalizedDocuments++;
				text = normalize(this.documents[i].text);
			}
			if (text.includes(needle)) ids.push(this.documents[i].id);
		}
		this.createdResultSlots += ids.length;
		if (!this.keepHistory) this.history.length = 0;
		this.history.push(ids);
		return ids;
	}

	inspect() {
		return {
			normalizedDocuments: this.normalizedDocuments,
			scannedDocuments: this.scannedDocuments,
			createdResultSlots: this.createdResultSlots,
			retainedResultSlots: this.history.reduce((sum, ids) => sum + ids.length, 0),
			retainedBatches: this.history.length,
			indexEntries: this.index?.length ?? 0
		};
	}
}


export function makeDocuments(count: number): Document[] {
	const topics = [
		'Factory service TypeScript',
		'Stack memory Go',
		'Graph service Rust',
		'Queue memory Go'
	];
	return Array.from({ length: count }, (_, id) => ({
		id,
		text: `  ${topics[id % topics.length]} lesson ${id}  `
	}));
}

export function makeQueries(count: number): string[] {
	return Array.from({ length: count }, (_, i) => queryCycle[i % queryCycle.length]);
}

export function replay(session: SearchSession, queries: string[]): number {
	let matches = 0;
	for (const query of queries) matches += session.search(query).length;
	return matches;
}

export function example(mode: Mode = 'baseline') {
	const session = new SearchSession(makeDocuments(1000), mode);
	const matches = replay(session, makeQueries(24));
	return { mode, matches, ...session.inspect() };
}

console.log(example());
GoComplete search workload
search.go
package main

import (
	"fmt"
	"strings"
)

type Document struct {
	ID   int    `json:"id"`
	Text string `json:"text"`
}

func normalize(text string) string { return strings.ToLower(strings.Trim(text, " \t\r\n")) }

type SearchSession struct {
	documents                                                 []Document
	index                                                     []string
	history                                                   [][]int
	keepHistory                                               bool
	normalizedDocuments, scannedDocuments, createdResultSlots int
}

func NewSession(documents []Document, mode string) *SearchSession {
	s := &SearchSession{documents: append([]Document(nil), documents...), keepHistory: mode == "baseline" || mode == "index"}
	if mode == "index" || mode == "both" {
		s.index = make([]string, len(documents))
		for i, document := range s.documents {
			s.index[i] = normalize(document.Text)
			s.normalizedDocuments++
		}
	}
	return s
}

func (s *SearchSession) Search(query string) []int {
	needle := normalize(query)
	ids := make([]int, 0)
	for i, document := range s.documents {
		s.scannedDocuments++
		var text string
		if s.index != nil {
			text = s.index[i]
		} else {
			text = normalize(document.Text)
			s.normalizedDocuments++
		}
		if strings.Contains(text, needle) {
			ids = append(ids, document.ID)
		}
	}
	s.createdResultSlots += len(ids)
	if !s.keepHistory {
		s.history = nil
	}
	s.history = append(s.history, ids)
	return ids
}


type Stats struct{ NormalizedDocuments, ScannedDocuments, CreatedResultSlots, RetainedResultSlots, RetainedBatches, IndexEntries int }

func (s *SearchSession) Inspect() Stats {
	retained := 0
	for _, ids := range s.history {
		retained += len(ids)
	}
	return Stats{s.normalizedDocuments, s.scannedDocuments, s.createdResultSlots, retained, len(s.history), len(s.index)}
}
func MakeDocuments(count int) []Document {
	topics := []string{"Factory service TypeScript", "Stack memory Go", "Graph service Rust", "Queue memory Go"}
	docs := make([]Document, count)
	for i := range docs {
		docs[i] = Document{i, fmt.Sprintf("  %s lesson %d  ", topics[i%len(topics)], i)}
	}
	return docs
}
func MakeQueries(count int) []string {
	cycle := []string{"FACTORY", " service ", "stack", "missing", "", "go", "memory", "factory"}
	queries := make([]string, count)
	for i := range queries {
		queries[i] = cycle[i%len(cycle)]
	}
	return queries
}

func Replay(s *SearchSession, queries []string) int {
	matches := 0
	for _, query := range queries {
		matches += len(s.Search(query))
	}
	return matches
}
func main() {
	for _, mode := range []string{"baseline", "index", "release", "both"} {
		s := NewSession(MakeDocuments(1000), mode)
		fmt.Println(mode, "matches:", Replay(s, MakeQueries(24)), "stats:", s.Inspect())
	}
}

The fixture is ASCII substring search with whitespace trimming and stable input order. It is not locale-aware search, ranking, networking, or a framework renderer. The runtime drivers use larger fixed workloads to collect samples; do not compare their elapsed times as a language ranking.

Leave this step with

An isolated change, preserved outputs, and an account of the work or lifetime it changed.

05 / Verify

Check the improvement. Keep the explanation.

Run the same behavior cases before trusting a timing. Then repeat the same workload with the profiler off for timing, and capture again to check whether the suspected hot path or retained objects changed.

The browser lab checks every query’s returned IDs, warms each version once, alternates order, and reports three runs including setup. This is a small local experiment. Three runs cannot establish a production percentile or a reliable speedup on every device. If the ranges overlap, say so and gather better evidence.

Before you call it fixed

  • Behavior: every query returns the same IDs in the same order, including empty and missing searches.
  • Resource: compare the same sample type and workload. Fewer total allocations in a shorter run is not a per-operation improvement.
  • Lifetime: repeat equivalent open → use → close cycles. Compare surviving state after cleanup; don't confuse heap size with process RSS.
  • Trade-off: include index construction, retained strings, and freshness. Measure the interaction or endpoint again where the symptom was reported.
Build UIs?See where this shows up in your components.

When the search is inside a frontend

If typing still stutters after filtering gets cheaper, inspect the rest of the interaction. A browser trace can reveal rendering, style, layout, painting, or other scripting work. The lab intentionally leaves these out. Optimizing a pure function does not establish that a page feels faster.

For memory, use a repeatable lifecycle: open a panel, interact, close it, then inspect what remains. A subscription, listener, timer, cache, or debug store can outlive the view. An allocation stack tells you where a result was created; a retaining path helps explain why it survives.

Allocation churn, sustained retention, and a large but stable working set are different investigations. Chrome’s memory guide covers these distinctions. Start from the user’s symptom, then choose the tool.

Field exercise

What would you investigate next?

Choose the evidence you need before choosing a fix.

A request takes 900 ms end to end. Its CPU samples account for little activity. You have not inspected the network or database spans.

Your next move
Your investigation note

Leave evidence the next engineer can follow.

What
The symptom, fixed workload, build, runtime, and machine.
Why
The profile observation and the cause it made you suspect.
Change
One intervention and the behavior it preserves or intentionally changes.
Evidence
Before/after captures, repeated timings, checks, and what remains uncertain.
Limits
Setup cost, lifetime, omitted work, and the next condition to test.

Try writing the conclusion without “it’s faster”: what stopped happening, what stopped being retained, and what did you pay for that change?

This technique connects measurement to design judgment. The ownership question will feel familiar from Composition over inheritance. The distinction between useful history and accidental retention connects to Stack’s undo history. Profiling supplies evidence for the decision; the application still decides which behavior matters.