← Home

Every "X Is Faster Than Y" Benchmark Has Three Knobs You Did Not See

Same code, same machine, opposite conclusions. Warmup, dead-code elimination and input size each decide the winner — and benchmark posts almost never say where they set them.

A stopwatch whose single hand has split into two, one pointing left and one pointing right.

Someone posts a benchmark. for beats reduce by 40%. The comments split into people who already believed it and people who run it themselves and get the opposite result. Both groups measured correctly.

The harness decides, and the harness is usually undocumented

A microbenchmark has at least three settings that change which implementation wins, and a benchmark post almost never states any of them:

  1. Warmup iterations. Engines interpret first and only compile to optimised machine code after a function has been called enough times. Zero warmup measures the interpreter plus the compiler — work your production code pays once, charged here to every iteration.
  2. Whether the result is used. Compute a value and throw it away and the engine may be able to prove the call has no observable effect, and delete it. You then measure an empty loop.
  3. Input size. Small inputs are dominated by call overhead, large ones by memory access. Implementations with identical asymptotic cost can swap places between those regimes.

Flip them yourself

Below are two implementations of exactly the same thing — summing an array. Run it at the defaults, then change one knob at a time.

Microbenchmark lab

Warmup iterations

Array size

Result

Results from your browser

Nothing measured yet.

The defaults are deliberately the bad configuration: zero warmup, result discarded. That is what a benchmark written in five minutes looks like, and it is the shape of most of the ones you will find.

I am not going to tell you which one wins, because on my machine the answer changed between settings and yours will too. That is the entire point — if a number moves when you change something the post did not mention, the number was never about the code.

What the three knobs actually do

Warmup. V8 runs a function in the interpreter, watches it, and only tiers it up to optimised code once it looks hot. With zero warmup you are timing the slow path plus the compilation. Production code that runs in a loop all day is always in the fast path, so a no-warmup benchmark is measuring a state your program is almost never in.

Dead-code elimination. This one produces the most spectacular results — code that appears to take almost no time, because it was removed. The fix is to make the result escape: accumulate it into a variable something later reads. In the component above that is exactly what the “accumulated” setting does, and it is one line.

Input size. This is the one people think they have handled by “using a realistic size”, but one size is one point. What you want is whether the lines cross, and you cannot see a crossing from a single measurement.

The honest use for microbenchmarks

They are good at rejecting, bad at choosing.

If a change is slower under every combination of settings, it is slower — that conclusion is robust. If it wins only at one warmup value and one input size, you have learned something about the harness, not about your program.

And when you do publish one, publish the knobs. Not because anyone will check, but because writing them down is what forces you to notice that you picked them.

Common follow-ups

Why does warmup change the answer so much?

JavaScript engines start by interpreting and only compile a function to optimised machine code after it has been called enough times. A benchmark with no warmup measures the interpreter, plus the cost of compiling — which is work your production code pays once and this benchmark charges to every iteration.

What is dead-code elimination doing to my benchmark?

If you compute a value and throw it away, the engine may be able to prove the whole call has no observable effect and delete it. You then measure an empty loop and conclude the code is free. Accumulating the result into something that is later read prevents the proof.

Does input size not just scale everything equally?

No. Small inputs are dominated by call overhead, and large ones by memory access. Two implementations with the same asymptotic cost can trade places between those regimes, so a single size is a single data point, not a curve.

So are microbenchmarks useless?

They are useful for rejecting things, not for choosing them. A change that is slower under every setting is genuinely slower. A change that wins only at one warmup value and one input size has told you nothing about production.