Methodology
Most IQ test sites publish nothing about how they score. This page is the opposite bet: every mechanism, every assumption, and every limitation of our tests, in plain language. Where we don't know something, it says so.
Every figural question is generated by rules, fresh for each attempt — not drawn from a stored question bank. Matrix puzzles are built by composing transformation rules (progression, row-constancy, Latin-square distribution) over shape, size, fill and count; sequences from six rule families; rotation figures as random chiral polycubes. Wrong answers are constructed by error type — the copy, the near-miss, the wrong-rule continuation — following the taxonomy used in the psychometric literature on matrix-item generation, because theory-typed distractors are what make an item measure reasoning rather than guessing style.
Before any question reaches you, an independent solver re-derives its answer from the visible content alone and verifies that exactly one option is defensible; sequences must have exactly one consistent continuation across all rule families, and every rotation target is verified chiral so its mirror image can never be a second correct answer. Our build fails if any generated item violates any of this — the checks run as ~7,000 automated assertions on every deploy.
The exception is the verbal domain: semantic relationships can't be safely machine-generated, so verbal analogies come from a hand-written bank with deliberately everyday vocabulary. It is the smallest domain, and the culture-fair test omits it entirely.
Scoring uses a two-parameter item response theory (IRT) model: each question carries a difficulty and a discrimination parameter, and your response pattern yields a probability distribution over ability (an EAP estimate). The centre of that distribution maps to the familiar IQ scale — mean 100, standard deviation 15 — and its spread is your error band. Getting hard questions right moves you more than easy ones; the same number correct can honestly produce slightly different estimates.
We report the band, not just the point, because the band is the measurement. On the 36-question test it is typically around ±5 points; on the 12-question quick test around ±12 — which is why that page tells you a short test can only place you in a broad band. We also clamp all scores to 55–145: outside ±3 standard deviations, a self-administered online test has no measurement claim to make, so we refuse to print numbers there.
Unanswered questions count as misses, and the instructions say so before you start — scoring and instructions never disagree.
An IQ score is a comparison against a reference population, so every IQ test lives or dies by its norm sample. Here is ours, stated plainly: v1 difficulty parameters are rational anchors — set from the structural complexity of each question type and the published behaviour of similar item families, not yet fitted to a population. As anonymous response data accumulates, the anchors are re-fitted and this page will state the sample size and refresh date.
What we will never claim: a representative national norm. People who take online IQ tests are a self-selected group — more educated and more interested in cognitive testing than average — and comparisons against such a pool inflate everyone's percentile. Clinical instruments solve this with stratified sampling census-matched on age, education, region and more; that is a large part of why a WAIS administration costs hundreds of dollars and this is free. If you see an online test claiming "normed on 500,000 test-takers" as if that were representative — quantity is not representativeness.
Timed administration measures a blend of reasoning and processing speed. The experimental literature on matrix tests finds time pressure costs roughly 3–5 raw points, with the penalty concentrated on older adults and test-anxious people — and psychometric comparisons of timed versus untimed short forms of Raven's Advanced Progressive Matrices found the untimed version the stronger instrument. Clinicians screening older adults use untimed matrices precisely so slower processing is not scored as weaker reasoning.
So these are power tests: the clock never runs, and pace cannot affect your score. We do record per-question response times silently — they feed question calibration and the descriptive "pace" note on your results — but no scoring pathway reads them. The one time-related interface element in the whole product is a gentle note if you idle for three minutes, reminding you that a best guess beats staring.
IQ is age-relative by definition: fluid reasoning peaks in the twenties and declines gradually while vocabulary holds until late life, so identical raw performance means different things at 25 and 70. Real instruments norm in age bands; online tests almost universally ignore age altogether.
Our position between those: children get separate tests per age tier (4–7, 8–12, 13–17) with their own difficulty calibration — never a re-graded adult test. Adults can optionally state an age band before starting; v1 scores everyone against a general adult reference and uses the band for context, because inventing adult age norms without data would be exactly the fake precision we criticise. The bands ride along with the anonymous response data, so real age-referenced scoring replaces the general reference as the sample grows — and this section will say when it has.
Children's results are graded by age in how much precision they claim: the 8–12 and 13–17 tiers report a wide, age-framed range — the width carries the uncertainty of provisional online norms — and the 4–7 tier reports descriptive bands only. A young child's point IQ without controlled age norms is precision theater, and we decline to print one.
We do not, and will not, norm by gender. Meta-analytic evidence finds no meaningful difference in full-scale IQ between men and women, so separate norms would be scientifically indefensible. But pretending the topic away would also be wrong, because subtest-level average differences are real and well replicated — 3D mental rotation shows the largest (favouring men on average, d ≈ 0.6); verbal fluency tasks tilt the other way. The design consequence is battery balance: spatial rotation is capped at about a fifth of every mixed test (7 of 36 on the full test), the domain weights are published, and no single gender-tilted format can dominate the score.
Anonymous response data also lets us run differential-item-functioning checks over time — the standard method for finding questions that behave differently across groups at the same ability level — and generator rules that produce such items get fixed or retired. Demographics are never requested before a test (being asked about group membership immediately before testing measurably shifts performance — stereotype threat) and are never required.
The culture-fair test extends the same principle to language and culture: no words anywhere, so vocabulary and schooling never move the score.
When you complete a test we store an anonymous response vector: which generated questions you saw (as a seed that can rebuild them), right/wrong per question, response times, and — for adults who chose to share it — an age band. No name, no email, no score, no free text. Adult submissions may carry the site's random anonymous session id; submissions from the kids, junior and teen tests carry no session id and no country, enforced server-side, so a minor's responses can never be joined to a browsing session or a location. This data exists for one purpose: re-fitting question difficulty so the test gets more accurate.
It cannot diagnose anything. It cannot qualify you for services, schools, or jobs — and employers should note that using cognitive scores in hiring carries specific legal obligations (job-relatedness and validation under U.S. employment law) that no online test satisfies. It is not a clinical instrument: a psychologist-administered WAIS-5 or Stanford–Binet, under controlled conditions against professionally gathered norms, remains the only route to a score that institutions should act on. And no IQ test of any kind measures creativity, judgment, persistence, or what you'll do with any of them.
What it can do, honestly: estimate your general reasoning ability within a stated band, show you how four reasoning domains compare, and be retaken meaningfully thanks to generated questions — for free, without taking your email or your time hostage.