What AsyncLocalStorage actually costs, measured

Sanjeev SharmaSanjeev Sharma
13 min read

Advertisement

Request-scoped context is the thing every service ends up needing: a request id in every log line, the current tenant in every query, a trace id that survives four awaits. AsyncLocalStorage does it without threading an argument through every function, and the objection is always the same — isn't that slow?

Here is the number, on three Node versions, with the script.

The short answer

Entering an AsyncLocalStorage context once per request costs about 250 nanoseconds on Node 24 and Node 26, and about 510 on Node 22. Each getStore() inside the context adds roughly 25 nanoseconds. At a thousand requests a second that is a quarter of a millisecond of CPU per second — about 0.025% of one core.

What was measured, exactly

A handler that awaits three times, which is what a real one does: a middleware hop, a database call, a serialisation step. Four variants of it, 200,000 iterations each, median of seven runs, inside official Alpine containers.

Per-request cost of request-scoped contextnanoseconds per request, lower is better
  • Node 22 · no context274
  • Node 22 · run() only784
  • Node 22 · run() + 3 reads825
  • Node 24 · no context250
  • Node 24 · run() only509
  • Node 24 · run() + 3 reads582
  • Node 26 · no context243
  • Node 26 · run() only492
  • Node 26 · run() + 3 reads570
  • Node 22 · run() only — the context itself costs 510 ns here
  • Node 24 · run() only — overhead halved against Node 22

Subtract the baselines and the picture is clean:

What you pay forNode 22Node 24Node 26
Entering a context (als.run)510 ns259 ns249 ns
Three getStore() calls inside it41 ns73 ns78 ns
One getStore() outside any context6 ns8 ns8 ns
Handler with no context at all (baseline)274 ns250 ns243 ns

The last row is the one people forget to measure. Three awaits cost about 250 nanoseconds on their own, so the question is never "is AsyncLocalStorage expensive" in isolation — it is "is it expensive next to the awaits you already have", and the answer is that it roughly doubles the cost of an empty handler and disappears entirely next to a handler that does work.

Why did it get cheaper in Node 24?

Because the propagation stopped going through an async_hooks callback on every async resource and started riding on V8's own continuation data. The design discussion is public — "AsyncLocalStorage without Async Hooks" lays out using v8::Context::SetContinuationPreservedEmbedderData() so the engine associates the frame with promise continuations directly, instead of Node observing every resource as it is created. The API surface in the async context docs did not change; what it costs did.

The measurement above is the observable half of that change — 510 ns down to 259 ns for the same code — and it lands between Node 22 and Node 24, which is where the engine moved from V8 12.4 to 13.6.

als.run({ requestId: "req_a" }, handler)getStore() → req_a at every hopcost: 249 ns to enter, 26 ns per readals.run({ requestId: "req_b" }, handler)interleaved on the same thread,never sees req_aoutside any run():getStore() → undefined, 8 nsthe alternative, threaded by hand:handler(req, ctx) → service(ctx) → repo(ctx) → query(ctx)zero runtime cost, and one forgotten parameter loses the trace

Two concurrent requests on one thread, each reading its own store. The cost is paid once at the boundary, not per await.

Is 250 nanoseconds per request ever a problem?

Only at a scale where you would already be counting nanoseconds. Put your own numbers in:

What request context costs your service

CPU spent on context, per second

0.56

(249 ns + 12 × 26 ns) × 1,000 req

Share of one core

0.056

context ms ÷ 1000 ms of wall clock

Added to a single request

0.0012

{(249 + reads * 26) / 1000|2} µs against 45 ms

Node 24 and 26 numbers. On Node 22, double the per-request figure. The reads slider is where teams surprise themselves: a logger that calls getStore() on every log line at 50 lines a request is 50 reads, not 1.

At a thousand requests a second with a dozen reads each, the whole mechanism costs about 0.06% of a core. The interesting failure is not the cost of the API — it is a logger that reads the store once per line and then formats a string for every line, which is a logging problem rather than a context problem.

When should you not use it?

OptionRuntime costRefactor costBreaks whenPick it when
AsyncLocalStoragedefault~250 ns/requestnone — wrap the entry pointcontext crosses a worker or a queueRequest-scoped values inside one process: request id, tenant, trace
Explicit parameter0every signature on the pathsomeone forgets to pass itA library, or a path short enough that the parameter is honest documentation
Module-level variable0nonethe second concurrent requestNever in a server. It is correct only in a single-request CLI
Re-reading from the request object0pass req everywherecode that is not in the request pathFramework middleware that already holds req and never leaves it
The runtime cost column is the least important one in this table, which is the point: a mechanism that is 250 ns and cannot be forgotten beats one that is free and can.

The real limit is the process boundary. Context does not travel to a worker_threads worker, into a queued job, or across an HTTP call — those need the value serialised and put back into a new context on the other side. Teams discover this when a background job's logs lose their request id, and the fix is to carry the id in the job payload rather than to reach for a bigger hammer.

How the measurement works

Reading the benchmark before trusting it
1 / 5

What this does not tell you

It does not measure memory. Every live context holds its store object alive for as long as the async tree under it is alive, so a store that accidentally captures a large response body keeps that body out of reach of the garbage collector for the life of the request. That is the failure mode worth watching, and it does not show up in a timing benchmark at all — it shows up as RSS that climbs under load, which the profiling walkthrough covers.

It also does not cover AsyncResource, which is what you need when you own a callback-based API and want context to survive it, or diagnostics channel, which is the better tool when the goal is instrumentation rather than request-scoped values. Both have their own costs and neither is measured here.

Check your model of the cost

3 questions — answers explained as you go.

  1. 1. Your handler does 12 getStore() calls on Node 26. What is the context cost per request?

  2. 2. A background worker logs without a request id. What is the fix?

  3. 3. Node 22 measured 510 ns of overhead and Node 24 measured 259 ns. Should that decide your upgrade?

Frequently Asked Questions

Is AsyncLocalStorage slow?

No. Measured on Node 24 and 26 it costs about 250 nanoseconds to enter a context per request and about 26 nanoseconds per getStore() call. On Node 22 the entry cost is roughly 510 nanoseconds. Both are negligible next to any handler that touches a network or a database.

Does AsyncLocalStorage leak memory?

The store stays reachable for as long as the async tree under the run() call is alive, which is correct behaviour and becomes a leak only if the store holds something large — a request body, a buffer, a database result. Keep stores to identifiers and small scalars.

Does context survive worker_threads or a job queue?

No. It is per-process and per-async-tree. Anything crossing a thread, a process or a network boundary has to carry the value explicitly and enter a new context on the other side.

Is it faster to pass a context object as a parameter?

Marginally, and it is the wrong trade for most services. A parameter costs nothing at runtime but has to be threaded through every signature on the path, and the failure mode is silent: one function that forgets it loses the trace for everything below it.

How do I measure this on my own runtime?

Run tools/bench/async-local-storage.mjs from this repository under your Node version. It prints the four cases and the per-request nanosecond figures, and the difference between the baseline and the run() case is your overhead.

The conclusion worth keeping

The cost of request context stopped being an argument two Node versions ago. If your service is slow, the cause is upstream of this: a query without an index, a serial chain of awaits that could have been parallel, or a response that serialises far more than the client reads — the three things the backend performance checklist puts first, and the ones the pagination post shows compounding.

If you are adding context to an existing service, start at the entry point in your Express setup, wrap the request there, and make sure your error handler reads the store before the stack unwinds — otherwise the one log line that needed the request id is the one that does not have it. The version-by-version picture of what else changed is in Node 22 vs 24 vs 26.

Advertisement

Sanjeev Sharma

Written by

Sanjeev Sharma

Full Stack Engineer · E-mopro

Related reading