Close scores need care.
The stability range reflects available evidence, not a guarantee of true ability. Small score gaps do not prove a meaningful advantage.
THE METHOD, SIMPLIFIED
We turn published AI benchmark results into one comparable score. So you can see how models stack up.
Explore the scoresSeven equally weighted domains
HOW IT WORKS
We collect published results and check their sources, model versions and test setups. We do not run the benchmarks ourselves.
See the evidenceEach benchmark counts once, using its best comparable reasoning setting. Related tests share influence; one evaluator’s influence is capped.
Seven domains have equal importance. A fixed scoring method combines them into one score, with a range showing its stability.
READING THE NUMBER
100 is the fixed reference average.The original reference models set the scale, with a standard deviation of 15. It is not the average of today’s models.
The scale can extend beyond these values.
KEEP IN MIND
The stability range reflects available evidence, not a guarantee of true ability. Small score gaps do not prove a meaningful advantage.
Missing test results stay missing. Unmeasured domains are estimated, and sparse or conflicting evidence can widen the range.
Price and speed stay separate. Results can favor models tested at more settings. Benchmarks do not guarantee performance on your work.