Worker threads: the break-even point, measured

Sanjeev SharmaSanjeev Sharma
13 min read

Advertisement

"Move it to a worker thread" is the standard advice for anything CPU-bound in Node, and it is wrong more often than it is right. Below a certain amount of work per task, the thread costs more than it saves — by a factor of thirty-five in the first row of this table.

Twelve cases, four payload sizes, three amounts of work each, one pool of four.

The short answer

With a pool of four workers, the break-even is around 2 ms of CPU per task. Below that, spawning and feeding workers costs more than doing the work inline. Above it, speedup climbs towards the core count — 2.6× at 14 ms a task, 3.5× at 222 ms. Transferring the payload instead of cloning it is what makes the large cases work at all.

The measured table

Each case runs 32 tasks three ways: inline on the main thread, across a pool of four with the payload structured-cloned, and across the same pool with the payload transferred. Work per task is a SHA-256 over the buffer, repeated 1, 8 or 64 times, which is a knob for CPU cost that does not change the payload size.

PayloadRoundsCPU per taskMain threadPool, clonedPool, transferredSpeedup
64 KB10.05 ms1.8 ms73.0 ms64.4 ms0.03×
64 KB80.26 ms8.2 ms55.0 ms61.8 ms0.13×
64 KB641.90 ms60.8 ms112.0 ms71.4 ms0.85×
512 KB10.22 ms7.2 ms67.3 ms49.8 ms0.14×
512 KB81.77 ms56.7 ms142.9 ms64.0 ms0.89×
512 KB6413.75 ms440.1 ms563.7 ms171.7 ms2.56×
2 MB10.86 ms27.5 ms99.3 ms75.9 ms0.36×
2 MB86.89 ms220.5 ms349.0 ms126.6 ms1.74×
2 MB6456.59 ms1,810.8 ms2,017.7 ms680.6 ms2.66×
8 MB13.44 ms110.2 ms358.3 ms175.3 ms0.63×
8 MB827.77 ms888.5 ms1,112.6 ms371.5 ms2.39×
8 MB64222.33 ms7,114.7 ms6,866.4 ms2,006.2 ms3.55×
Same 32 tasks, three ways — 8 MB payloadtotal milliseconds for the batch, lower is better
  • 1 round · main thread110.2
  • 1 round · pool, transferred175.3
  • 1 round · pool, cloned358.3
  • 64 rounds · main thread7,115
  • 64 rounds · pool, cloned6,866
  • 64 rounds · pool, transferred2,006
  • 1 round · pool, transferred — the pool loses: 3.4 ms of work per task is not enough
  • 64 rounds · pool, cloned — copying 8 MB per task eats most of the parallelism
  • 64 rounds · pool, transferred — 3.55× — near the 4× ceiling of a four-worker pool

Where exactly is the line?

Read down the speedup column and it crosses 1.0 between 1.77 ms and 13.75 ms of CPU per task. The two cases either side of it are the interesting ones: at 1.90 ms the pool is still 15% slower, and at 13.75 ms it is 2.6× faster. So the practical rule is a few milliseconds of CPU per task, and below that the answer is to do the work inline.

Should this task go to a worker?

Estimated speedup with transfer

2.40×

(12 ms × 32) ÷ ((12 + overhead) × 32/4 + 60 ms pool start)

Verdict

worth a pool

a pool that wins by less than 30% is rarely worth the failure modes

Break-even CPU per task

2.66 ms

overhead spread across 4 workers and 32 tasks

Fitted to the twelve measured cases above: roughly 60 ms to start a four-worker pool, 0.35 ms of scheduling per task, and transfer cost proportional to payload. It is an estimate — the script prints your real numbers in about a minute.

Why is cloning so much worse than transferring?

Because a structured clone copies the bytes and a transfer moves the ownership. At 8 MB per task and 32 tasks, that is 256 MB copied into the worker and nothing copied back — and the table shows exactly what it costs: 6,866 ms cloned against 2,006 ms transferred, for identical work.

What happens to your buffer on postMessage
1 / 5

The benchmark measures clone and transfer separately for exactly this reason: most worker-thread disappointment is a copy nobody intended.

The transferable list and its semantics are specified in the worker_threads documentation, and the underlying rules come from the HTML structured clone algorithm that Node reuses.

What the table does not include

Pool lifetime. Every case here creates four workers and terminates them, which is what a request-scoped pool does and what the "spawn a worker for this upload" pattern costs. That start-up shows up plainly in the smallest case: 1.8 ms of actual work took 64.4 ms wall clock, so roughly 60 ms went on creating and tearing down the pool — about 15 ms per worker.

A long-lived pool pays that once at boot instead of once per batch, which moves the break-even down but does not remove it: the per-task scheduling and transfer costs stay. If your service already keeps a pool warm, subtract the 60 ms and re-read the table; the small cases are still losses.

1.0× — the pool breaks even here4.0×1.0×0.0×0.050.260.861.913.827.856.6222CPU per task, milliseconds (not to scale)below the line, inline wins

Transferred-payload speedup from the table. The curve flattens towards 4× because the pool has four workers — no amount of extra work per task buys more parallelism than there are cores.

When is a worker the right answer anyway?

OptionCPU per taskPayloadMeasured outcomePick it when
Image resize, PDF render, video thumbnaildefault50–500 msMBs, transferable2.4–3.5× with transferThe work is long, the payload is a buffer you can hand over
Crypto over large blobs10–200 msMBs2.5× and upHashing or encrypting anything above a few hundred kilobytes
JSON parse of a big request1–10 msthe string itself0.1–0.9× — usually a lossRarely. The string has to be cloned into the worker, which costs what the parse cost
Template rendering, validation, mappingunder 1 mssmall objects0.03–0.4×Never. This is the case that made worker threads look bad
Blocking the loop on purposeanyanyn/aA separate process, not a thread, when the work can crash or run unbounded
The column that decides it is the first one. Payload size changes how much of the win you keep; CPU per task decides whether there is a win to keep.

Two failure modes are worth naming. The first is the one in row four: a service that moves per-request JSON parsing into a pool and gets slower, because the string is cloned into the worker and the parse was never the bottleneck. The second is unbounded pools — one worker per request, which turns a CPU problem into a memory problem at about 15 ms and several megabytes per worker.

If the work can hang or crash, a worker is also the wrong isolation boundary: a thread shares the process, so an out-of-memory in the worker takes the whole service with it. That is what child processes are for — a separate heap, a separate crash, and a SIGKILL you can actually use.

Would you reach for a worker here?

3 questions — answers explained as you go.

  1. 1. Your handler parses a 512 KB JSON body in about 1.8 ms. Move it to a pool?

  2. 2. You resize 32 images, 8 MB each, about 220 ms of CPU apiece. Clone or transfer?

  3. 3. A four-worker pool gives 3.55× on the heaviest case. Why not 4×?

Frequently Asked Questions

When are worker threads worth it in Node?

When one task needs more than a few milliseconds of CPU. Measured with a pool of four: below about 2 ms per task the pool is slower than the main thread, at 13.75 ms it is 2.6× faster, and at 222 ms it reaches 3.55×.

Why is my worker pool slower than doing the work inline?

Almost always because the task is too small or the payload is being cloned. Starting four workers costs roughly 60 ms, and a structured clone copies the whole payload on the sending thread — at 8 MB per task that is more expensive than the work itself.

What is the difference between transferring and cloning a buffer?

A clone copies the bytes into the receiving thread and leaves yours intact. A transfer detaches the ArrayBuffer here and reattaches it there, moving zero bytes — measured at 2,006 ms against 6,866 ms for the same batch.

How many workers should a pool have?

Start at availableParallelism() and no more. Speedup is capped by cores, so extra workers add memory and scheduling without adding throughput — the measured 3.55× on four CPUs is already close to the ceiling.

Should I use worker threads or child processes?

Threads for CPU work you trust, inside your own process and memory. Processes when the work can crash, hang or run unbounded, because a worker thread that runs out of memory takes the whole service down with it.

The rule to keep

Measure one task before you build a pool. If it takes less than a couple of milliseconds of CPU, the pool will lose, and no amount of tuning changes that — the overhead is a fixed tax and small tasks cannot pay it.

The same discipline applies to everything else that gets blamed for slow responses: the backend performance checklist puts measurement before architecture for the same reason, and profiling usually finds that the event loop was blocked by something far cheaper to fix — a regex with nested quantifiers, a synchronous log write, or a query without an index.

For the runtime around it: what changed between Node versions covers the serialiser work that makes big responses cheaper, request context costs about 250 nanoseconds, and running TypeScript directly removes a build step from the loop you are about to profile.

Advertisement

Sanjeev Sharma

Written by

Sanjeev Sharma

Full Stack Engineer · E-mopro

Related reading