← Concepts & practices
Pattern Concurrency, scheduling, and delivery

Bounded parallelism

How much work can be in flight?

A batch of 400 independent calls looks like an opportunity to go faster. Starting all 400 at once also turns your process, database, or provider into the queue. Bounded parallelism makes the admission rule explicit: choose how many jobs may be active, what waits, and what happens when the queue itself is full.

The judgment to keep

A concurrency bound protects a resource only when its number comes from that resource’s budget. Sequential work, a p-limit-style limiter, and a worker pool all make different queue and shutdown promises. Unbounded Promise.all is not neutral; it admits the whole input at once.

TypeScriptGo One job list · four admission policies
Start with the pressure

Faster for one batch can be slower for everyone.

Suppose an import has 400 independent records. A sequential loop keeps the dependency calm but adds every latency together. Promise.all(records.map(send)) finishes quickly when the dependency has spare capacity, but it also submits all 400 calls immediately.

The useful middle is to admit a known number of jobs. A limiter keeps a finite batch in memory and starts only the next job when a slot opens. A worker pool names long-lived consumers and can make the queue itself bounded, so a continuous producer must wait, shed, or retry instead of growing memory forever.

“Concurrent” tells you that work overlaps. It does not tell you how much overlap your dependency can survive.

Read the safe baselineTypeScript · one job at a time
bounded.ts · sequential baseline
export async function runSequential(jobs: Job[], work: Work): Promise<string[]> {
	const results: string[] = [];
	for (const job of jobs) results.push(await work(job));
	return results;
}

Sequential work is not automatically the right production choice, but it is a useful baseline: maximum active work is one, ordering is obvious, and no hidden queue can grow. Every faster policy should be able to explain what it gives up.

Name the admission policy

Four ways to decide what runs next.

The policies differ in more than syntax. Sequential work has no queue. Unbounded fan-out delegates capacity to the runtime and dependency. A p-limit-style helper caps active Promises but can retain a huge pending list. A worker pool caps consumers and gives a long-lived queue a place to apply backpressure and shutdown.

Policy 01

Sequential

One current job, no overlap, the simplest ordering and cleanup story.

Active
Exactly one.
Queue
Only the caller’s next item.
Best for
Order, shared state, or tiny batches.
Policy 02

Unbounded

Start every job immediately and let the dependency absorb the burst.

Active
Every submitted item.
Queue
Hidden in promises, sockets, or the provider.
Best for
Only inputs already known to be small.
Policy 03

p-limit style

Keep a finite Promise batch, but start only up to the configured limit.

Active
The limit.
Queue
Pending jobs in memory.
Best for
Finite async work with a known cap.
Policy 04

Worker pool

Fixed consumers pull jobs from a queue whose capacity and shutdown are explicit.

Active
The worker count.
Queue
A channel or queue with an admission policy.
Best for
Continuous work and named worker lifetimes.
What each policy actually bounds
PolicyActive workPending workWhat it still needs
Sequential1Caller-controlledEnough total time
UnboundedAll submittedHidden or dependency-ownedA trusted input and capacity budget
p-limitConfigured limitIn-memory listQueue size, cancellation, retry policy
Worker poolWorker countExplicit queueClose, shutdown, and queue-full behavior
TypeScript · p-limit style
bounded.ts · p-limit style
export function runUnbounded(jobs: Job[], work: Work): Promise<string[]> {
	return Promise.all(jobs.map(work));
}

type Pending = {
	run: () => Promise<unknown>;
	resolve: (value: unknown) => void;
	reject: (reason: unknown) => void;
};

// A small p-limit-style helper: active work is capped, while pending work waits in memory.
export function createLimiter(concurrency: number) {
	if (!Number.isInteger(concurrency) || concurrency < 1) {
		throw new RangeError('concurrency must be a positive integer');
	}
	let active = 0;
	const pending: Pending[] = [];

	function pump() {
		while (active < concurrency && pending.length > 0) {
			const item = pending.shift()!;
			active += 1;
			void item
				.run()
				.then(item.resolve, item.reject)
				.finally(() => {
					active -= 1;
					pump();
				});
		}
	}

	return function limit<T>(run: () => Promise<T>): Promise<T> {
		return new Promise<T>((resolve, reject) => {
			pending.push({
				run,
				resolve: (value) => resolve(value as T),
				reject
			});
			pump();
		});
	};
}

export async function runWithLimit(
	jobs: Job[],
	concurrency: number,
	work: Work
): Promise<string[]> {
	const limit = createLimiter(concurrency);
	return Promise.all(jobs.map((job) => limit(() => work(job))));
}

// A worker-pool style loop keeps a fixed number of consumers pulling from one job list.
export async function runWorkerPool(jobs: Job[], workers: number, work: Work): Promise<string[]> {
	if (!Number.isInteger(workers) || workers < 1) {
		throw new RangeError('workers must be a positive integer');
	}
	const results = new Array<string>(jobs.length);
	let next = 0;
	async function worker() {
		while (next < jobs.length) {
			const index = next;
			next += 1;
			results[index] = await work(jobs[index]);
		}
	}
	await Promise.all(Array.from({ length: Math.min(workers, jobs.length) }, worker));
	return results;
}
Go · semaphore and worker pool
bounded.go · semaphore and worker pool
// One goroutine per job: every job is admitted at once.
func runUnbounded(ctx context.Context, jobs []Job, work Work) ([]string, error) {
	results := make([]string, len(jobs))
	errs := make([]error, len(jobs))
	var wg sync.WaitGroup
	for index, job := range jobs {
		wg.Add(1)
		go func() {
			defer wg.Done()
			results[index], errs[index] = work(ctx, job)
		}()
	}
	wg.Wait()
	return results, errors.Join(errs...)
}

// A p-limit-style semaphore: a goroutine per job, but only `limit` may hold a slot.
func runWithLimit(ctx context.Context, jobs []Job, limit int, work Work) ([]string, error) {
	if limit < 1 {
		return nil, fmt.Errorf("limit must be positive")
	}
	slots := make(chan struct{}, limit)
	results := make([]string, len(jobs))
	errs := make([]error, len(jobs))
	var wg sync.WaitGroup
	for index, job := range jobs {
		wg.Add(1)
		go func() {
			defer wg.Done()
			slots <- struct{}{}        // wait for a free slot
			defer func() { <-slots }() // give it back
			results[index], errs[index] = work(ctx, job)
		}()
	}
	wg.Wait()
	return results, errors.Join(errs...)
}

// A fixed number of workers pull job indexes from an explicit, bounded queue.
func runWorkerPool(ctx context.Context, jobs []Job, workers int, work Work) ([]string, error) {
	if workers < 1 {
		return nil, fmt.Errorf("workers must be positive")
	}
	queue := make(chan int, workers) // a full queue makes the producer wait
	results := make([]string, len(jobs))
	errs := make([]error, len(jobs))
	var wg sync.WaitGroup
	for range workers {
		wg.Add(1)
		go func() {
			defer wg.Done()
			for index := range queue {
				results[index], errs[index] = work(ctx, jobs[index])
			}
		}()
	}
	for index := range jobs {
		queue <- index
	}
	close(queue)
	wg.Wait()
	return results, errors.Join(errs...)
}
Follow the queue

Change admission. Watch pressure move.

Choose a policy and a workload. The lab keeps the job list conceptually fixed so you can ask what the policy changes: active count, pending queue, dependency pressure, and the work needed to shut down cleanly.

One job queue

Change admission. Watch pressure and overlap.

Runs a local comparison
Unbounded Promise.all Saturated dependency
01map every job
02all start together
03dependency absorbs the burst
Max in flight

Every submitted job

Queue

The runtime and dependency absorb the whole batch at once.

Overlap

Maximum possible overlap; actual parallelism depends on the runtime and resource.

Pressure

A burst can exhaust sockets, rate limits, memory, or downstream capacity.

Only use when the input and dependency budget are already bounded.

Watch for Promise.all joins the batch; it does not bound admission.

The controls change a local model; they do not start real workers or network calls.
Read both implementationsTypeScript and Go · same job list, metered

Sequential baseline: one job finishes before the next starts.

TypeScriptReading
bounded.ts
export async function runSequential(jobs: Job[], work: Work): Promise<string[]> {
	const results: string[] = [];
	for (const job of jobs) results.push(await work(job));
	return results;
}
GoAlongside
bounded.go
func runSequential(ctx context.Context, jobs []Job, work Work) ([]string, error) {
	results := make([]string, 0, len(jobs))
	for _, job := range jobs {
		value, err := work(ctx, job)
		if err != nil {
			return nil, err
		}
		results = append(results, value)
	}
	return results, nil
}
Name what is bounded

Choose the smallest honest admission contract.

Practice distinguishing active concurrency from queue memory, throughput from parallelism, and a finite batch from a producer that never stops. The most useful answer names the resource and what happens when its budget is exhausted.

A finite import has 400 independent HTTP calls and the provider allows 8 in flight.
A never-ending producer can create jobs faster than the dependency completes them.
A Go service needs to process jobs continuously and shut down cleanly.
Each task mutates one shared in-memory accumulator and must observe input order.
Feedback stays on this page; it is not saved.
A production boundary

Put the limit beside the budget it protects.

For a finite batch of remote calls, a limiter can be the smallest useful tool. Keep the input and result order explicit, make failures observable, and choose whether a canceled batch drains, drops pending work, or waits for active requests to settle. Do not confuse a limit of eight with a provider’s eight-per-second rate limit; those are different dimensions.

For continuous work, move from a list of pending Promises to a queue with a capacity policy. A worker pool can block the producer, reject new jobs, or shed the oldest work. Its shutdown path must stop admission, close the queue, let workers finish or cancel, and then join them.

Budget

Which resource is scarce?

Set the number from sockets, CPU, memory, rate, or latency.

Admission

What waits or gets refused?

Make active work and queue capacity visible to the producer.

Shutdown

Who closes the workers?

Stop intake, settle or cancel jobs, and join every consumer.

A card loader admits two requests at a time and aborts the effect-owned work on cleanup.

ReactAlready in your code
textbook.tsx · two-at-a-time loader
import { useEffect, useState } from 'react';

type Card = { id: string; value: unknown };

async function loadCard(id: string, signal: AbortSignal): Promise<Card> {
	const response = await fetch(`/api/cards/${id}`, { signal });
	if (!response.ok) throw new Error(`${response.status} from ${id}`);
	return { id, value: await response.json() };
}

async function mapWithLimit<T, R>(items: T[], limit: number, work: (item: T) => Promise<R>) {
	const results = new Array<R>(items.length);
	let next = 0;
	async function worker() {
		while (next < items.length) {
			const index = next++;
			results[index] = await work(items[index]);
		}
	}
	await Promise.all(Array.from({ length: Math.min(limit, items.length) }, worker));
	return results;
}

export function AccountCards({ ids }: { ids: string[] }) {
	const [cards, setCards] = useState<Card[]>([]);
	const [error, setError] = useState<string | null>(null);

	useEffect(() => {
		const controller = new AbortController();
		let current = true;
		void mapWithLimit(ids, 2, (id) => loadCard(id, controller.signal)).then(
			(next) => current && setCards(next),
			(reason) => {
				if (current && !controller.signal.aborted) {
					setError(reason instanceof Error ? reason.message : String(reason));
				}
			}
		);
		return () => {
			current = false;
			controller.abort();
		};
	}, [ids]);

	return error ? (
		<p>{error}</p>
	) : (
		<ul>
			{cards.map((card) => (
				<li key={card.id}>{card.id}</li>
			))}
		</ul>
	);
}
Build services or UIs?The capacity budget is already part of the feature contract.

Where it already is in your components

A data loader’s batch size, a browser’s simultaneous image fetches, and a Go jobs channel all answer “how many can be active?” Review the queue and shutdown answers beside that number.

When you have to own it

When a dependency has a budget or a producer can outpace consumers, own the admission policy at the boundary that knows the cost. Keep queue age, rejected work, and active count observable.

Recognize it in UI code

map is often an admission decision in disguise.

Finite batch

A limiter can keep all job identities in memory and restore results by index.

Growing list

Use a queue or window; a pending Promise for every item is still memory pressure.

Changing view

Cancel the old scope before starting new work so the limit protects the current screen.

Pair the bound with a structured scope ↗
The parts to watch

The number alone does not make a system safe.

p-limit does not bound the pending list

A limiter can leave thousands of closures and inputs waiting in memory. For unbounded producers, add a bounded queue and a policy for full admission. The active count and queue count are two different measurements.

Concurrency is not a rate limit

Eight active requests can still produce more than eight calls per second if they complete quickly. Use a token bucket, delay, or provider-aware scheduler when the contract is temporal.

A worker pool needs a close protocol

Decide who stops producers, who closes the jobs channel, whether queued jobs drain, how workers observe cancellation, and who joins them. Without that protocol, a fixed count can still leak workers or strand sends.

More workers can reduce throughput

Contention, connection pools, garbage collection, throttling, and cache misses can make a larger limit slower. Treat the limit as a measured hypothesis and watch dependency latency, errors, and queue age.

Make the call

Choose the bound that matches the bottleneck.

Use sequential work when order or shared mutation is the constraint. Use unbounded fan-out only when both input size and dependency capacity are trusted. Use a p-limit-style helper for a finite batch when active work is the only bound you need. Use a worker pool when work is continuous or when queue, shutdown, and worker lifetime deserve explicit names.

One at a time

Keep the order obvious.

Safe baseline for shared state or sequential dependencies.

Finite batch

Limit active work.

Use a dependency-backed cap and observe queue memory.

Continuous flow

Name the pool and queue.

Make backpressure, shutdown, and rejection part of the protocol.

Take the idea with you

Make “how many?” a real design question.

Unbounded fan-out is easy to write because the input list silently becomes the concurrency setting. Bounded parallelism breaks that shortcut apart: how many are active, how many wait, who admits them, and what happens when the budget is gone? The limit protects the resource; the owner closes the conversation.

Why
Protect dependency capacity and keep progress observable.
What
Bound active work and name the pending-work policy.
Constraint
The cap must preserve correctness, cancellation, and shutdown, not only throughput.
Fallback
Start with a smaller bound. It may lower throughput, so measure before increasing it.
Reconsider when
The real constraint is a rate, priority, or durable queue.
Connections to follow nextRelated lessons
Explore more concepts & practices →