The whole recipe.
There is no other one.
HX is a self-administered benchmark. It measures how you perform on eight short tests — against every other human who has taken them, and against frontier AI models run through the same events under the same constraints.
The baselines
4 models · run 2026-08Every machine score on this site comes from a versioned baseline run: each frontier model is run through the exact same eight events a human plays, under the same time limits, with the same shuffled pools. No cherry-picking across attempts — a baseline is one complete run protocol, executed three times, median kept.
We baseline several models — multiple labs, plus a strong open model. The scoreboard compares you against the champion per dimension: a dimension is held only if the human median beats the best model at it. Baselines are stamped wherever they appear. When a new frontier model ships, we re-run the full protocol and a new season begins. Old baselines stay published.
The human side
pre-launch honestystillhuman launches as Season 01. There are no prior seasons and we won't pretend there were.
Until enough humans have played, the human reference values on the scoreboard come from published literature — reaction-time distributions, digit-span norms, detection-accuracy studies — each traceable to its source. They are labeled “human reference · pre-launch” wherever shown.
Once a dimension reaches 5,000 scored human runs, the reference is retired and the live human median takes over. The scoreboard will say which one you're looking at. Always.
the “judgment fell nov 2025” date refers to when published model evaluations first cleanly exceeded the human reference on our judgment protocol — a baseline event, not a stillhuman season.
Percentiles
provisional until n ≥ 5,000Your percentile places you among every human who has completed the same event this season. Below 5,000 scored runs per event, percentiles carry a “Provisional” flag and may shift as the pool grows. We show the flag instead of hiding the uncertainty.
Percentiles reset each season. Your card keeps its season stamp, so an 83rd percentile from the GPT-5 era stays exactly what it was.
Held, lost, and margins
the scoreboard rulesA dimension is “held” when the human median beats the best machine baseline on that dimension's events, “lost” when it doesn't. The landing page shows only these two states. Dimensions within ±3 points are internally marked contested and re-run monthly; the scoreboard flips only on two consecutive re-runs agreeing.
Margins update every season, or sooner if a baseline re-run is triggered by a major model release mid-season.
Timing and anti-cheat
server-timedAll reaction timing is server-validated; client timestamps are cross-checked and implausible runs are voided, not penalized. False starts void the round. Item pools (300+ per perception event) shuffle every run, so there is nothing to study and no retake-grinding worth the effort.
One scored run per event per day counts toward your card. Practice runs are unlimited and unscored.
What HX is not
constraintsHX is a self-administered benchmark taken in a browser, on your hardware, at your hour of the day. It is not a hiring assessment, not an IQ test, and not a clinical instrument. It measures performance on these eight events under these rules — nothing broader, and we won't claim otherwise.