mechanix.gg

What Makes a Good Mechanics Benchmark?

Most reaction tests measure the wrong thing, or measure it badly. The six properties a benchmark needs before its score means anything, and how we built ours.

Piotr VassevPiotr Vassev··3 min read
A caster stepping out of a bush and casting a sigil at the archer

A benchmark is only as good as the thing it measures and how carefully it measures it. Most online reaction tests do one of these well. Here's the checklist we hold ours to.

1. It measures the skill you actually use

Clicking when a box turns green measures simple reaction time. Games demand recognition, decision and accurate execution under pressure. A good mechanics benchmark recreates that:

  • Realistic input: right-click movement, an aimed Flash, an ability key, not one big button.
  • Realistic threats: a charge with a landing zone, a hook that leads you, a stun you have to cleanse.
  • An outcome, not just a time: did you escape, dodge or cleanse? A fast reaction in the wrong direction is still a failure.

2. The timing is honest

Small timing errors are the same size as the differences you're trying to measure. So:

  • Time from what you can see. We start the clock at the first rendered frame that shows the threat, not when the game logic decided it would happen.
  • Use your input's own timestamp. Keyboard and mouse events carry the time they happened, so the measurement doesn't depend on when the next frame happened to process them.
  • Simulate in fine steps. Our games run on a fixed 1 ms simulation step, with each input applied at its own time, so results don't depend on your frame rate.
  • Throw out broken attempts. If a frame takes longer than 50 ms during the measured window, the timing can't be trusted, so we repeat that attempt.

3. You can't game it with timing

If the threat comes at a predictable moment, people learn the rhythm and press early. Good benchmarks:

  • Randomize the wait. Our cues come after a random 1.5–5 s delay.
  • Randomize the threat. Distances, angles and projectile speeds vary, so you can't memorize the landing time.
  • Reject anticipation. Presses under ~100 ms after the cue don't count as reactions. Cleanse is the deliberate exception: it measures timing, so pressing before the stun is a wasted Cleanse and a failed attempt.

4. Enough attempts, summarized robustly

One attempt is noise. Every session has 10 scored attempts, summarized by the median, which resists outliers, and a consistency figure, the spread of your reactions. A player with a 210 ms median and ±15 ms spread is more reliable than one with a 190 ms best and ±80 ms spread.

5. The score is transparent

A single number is convenient but can hide a lot. Ours is a 0–1000 blend of median reaction, success rate and consistency (plus direction where it applies), with the weights shown on the results page, and every attempt logged with its outcome, timings and frame stability. If you disagree with the weighting, the raw numbers are all there.

6. It's reproducible

Every session runs from a seed. Share a challenge link and your friend gets the same scenario: same positions, same timings, same threats. That makes comparisons fair and lets you retest yourself on identical conditions.

Try it · free, in your browserCleanseA homing sigil always hits. Cleanse the stun within 100 ms.

What no browser benchmark can fix

Honesty includes limits. Your hardware adds delay we can't fully see: display scanout, panel response and input device lag. We record your refresh rate and flag unstable frames, but a 60 Hz office monitor and a 240 Hz gaming display will never be identical. Compare yourself on the same setup, and compare others with that in mind.

A good benchmark is one where a better score means you got better at the game skill, not at the test. That's the bar we build to.

FAQ

Why do reaction tests give different results?
They differ in what they measure, a simple click versus a recognition-and-response task, and in how they time it. Hardware delay, frame timing, and whether anticipation is rejected all change the number.
How many attempts do you need for a reliable reaction time?
Single attempts are noisy. Ten attempts summarized by the median gives a much more stable picture than a best-of or an average that one outlier can drag around.
Why does mechanix.gg reject some inputs as too early?
A press faster than about 100 ms after a cue is almost certainly a guess, not a reaction. Counting guesses would reward anticipation instead of reaction. Cleanse is the exception: it's deliberately a timing test, so pressing before the stun simply fails the attempt.
Piotr Vassev

Piotr Vassev

Software engineer, long-time League player and creator of mechanix.gg. Builds the frame-accurate timing and game simulations behind the benchmarks, and writes about what the numbers actually say.

Connect on LinkedIn

Keep reading