THE METHOD, SIMPLIFIED

Many benchmarks.
One clearer picture.

We turn published AI benchmark results into one comparable score. So you can see how models stack up.

Explore the scores
FROM EVIDENCE TO SCORE
92benchmarks in the method

Seven equally weighted domains

1The Latent Score
Higher means stronger benchmark performance.

HOW IT WORKS

From results to a ranking.

  1. 01

    Start with evidence.

    We collect published results and check their sources, model versions and test setups. We do not run the benchmarks ourselves.

    See the evidence
  2. 02

    Make results comparable.

    Each benchmark counts once, using its best comparable reasoning setting. Related tests share influence; one evaluator’s influence is capped.

  3. 03

    Bring it all together.

    Seven domains have equal importance. A fixed scoring method combines them into one score, with a range showing its stability.

THE SEVEN DOMAINS
  • Reasoning
  • Coding
  • Tool use
  • Professional work
  • Vision
  • Cybersecurity
  • Communication

READING THE NUMBER

A reference point.
Not an IQ test.

100 is the fixed reference average.The original reference models set the scale, with a standard deviation of 15. It is not the average of today’s models.

THE SCORE SCALEHigher is better →
7085100Reference average115130

The scale can extend beyond these values.

KEEP IN MIND

Useful context. Honest limits.

Close scores need care.

The stability range reflects available evidence, not a guarantee of true ability. Small score gaps do not prove a meaningful advantage.

Missing is not zero.

Missing test results stay missing. Unmeasured domains are estimated, and sparse or conflicting evidence can widen the range.

Capability is one part.

Price and speed stay separate. Results can favor models tested at more settings. Benchmarks do not guarantee performance on your work.