methodology · v1.0season 01 · gpt-5 erathis page is the whole recipenothing quietly replaced
published · aug 2026
how the numbers are made

The whole recipe.
There is no other one.

HX is a self-administered benchmark. It measures how you perform on eight short tests — against every other human who has taken them, and against frontier AI models run through the same events under the same constraints.

fig. m1 · how a score is made~6 minutes, start to card
01you play
02server validates
03percentile vs humans
04verdict vs baseline
05card updates

The baselines

4 models · run 2026-08

Every machine score on this site comes from a versioned baseline run: each frontier model is run through the exact same eight events a human plays, under the same time limits, with the same shuffled pools. No cherry-picking across attempts — a baseline is one complete run protocol, executed three times, median kept.

We baseline several models — multiple labs, plus a strong open model. The scoreboard compares you against the champion per dimension: a dimension is held only if the human median beats the best model at it. Baselines are stamped wherever they appear. When a new frontier model ships, we re-run the full protocol and a new season begins. Old baselines stay published.

baselinerunprotocolkept
gpt-52026-083× fullmedian · live
claude 4.52026-083× fullmedian · live
gemini 32026-083× fullmedian · live
next frontier (?)———3× fullawaiting release

The human side

pre-launch honesty

stillhuman launches as Season 01. There are no prior seasons and we won't pretend there were.

Until enough humans have played, the human reference values on the scoreboard come from published literature — reaction-time distributions, digit-span norms, detection-accuracy studies — each traceable to its source. They are labeled “human reference · pre-launch” wherever shown.

Once a dimension reaches 5,000 scored human runs, the reference is retired and the live human median takes over. The scoreboard will say which one you're looking at. Always.

the “judgment fell nov 2025” date refers to when published model evaluations first cleanly exceeded the human reference on our judgment protocol — a baseline event, not a stillhuman season.

Percentiles

provisional until n ≥ 5,000

Your percentile places you among every human who has completed the same event this season. Below 5,000 scored runs per event, percentiles carry a “Provisional” flag and may shift as the pool grows. We show the flag instead of hiding the uncertainty.

Percentiles reset each season. Your card keeps its season stamp, so an 83rd percentile from the GPT-5 era stays exactly what it was.

Held, lost, and margins

the scoreboard rules

A dimension is “held” when the human median beats the best machine baseline on that dimension's events, “lost” when it doesn't. The landing page shows only these two states. Dimensions within ±3 points are internally marked contested and re-run monthly; the scoreboard flips only on two consecutive re-runs agreeing.

Margins update every season, or sooner if a baseline re-run is triggered by a major model release mid-season.

Timing and anti-cheat

server-timed

All reaction timing is server-validated; client timestamps are cross-checked and implausible runs are voided, not penalized. False starts void the round. Item pools (300+ per perception event) shuffle every run, so there is nothing to study and no retake-grinding worth the effort.

One scored run per event per day counts toward your card. Practice runs are unlimited and unscored.

What HX is not

constraints

HX is a self-administered benchmark taken in a browser, on your hardware, at your hour of the day. It is not a hiring assessment, not an IQ test, and not a clinical instrument. It measures performance on these eight events under these rules — nothing broader, and we won't claim otherwise.

change log
v1.0 · aug 2026 · initial publication · baselines: 4 frontier models · run 2026-08nothing has been quietly replaced. this log is where we'd have to admit it.
hx is not a hiring assessment · not an iq test · not science we don't havethe scoreboard →