The comfortable number
AI leaderboards often give you one number. A model scores 80%, and the number reads like a verdict even though it is an average.
Averaging smooths over variance. An 80% aggregate could come from similar results in every category or from perfect scores in eight equally weighted categories and zero in two.
The two failed categories still matter. The mean makes them harder to see.
What the average buries
Picture a compliance benchmark spanning ten equally weighted frameworks. The model scores well on GDPR, ISO 27001, and SOC 2, but poorly on the EU AI Act, the Data Act, and DORA.
A high average could still suggest broad reliability even when the model is unsuitable for some regimes. The headline score does not show whether it can handle the framework a buyer needs.
Composite scores compress many results into one ranking. Research on benchmark aggregation notes that central leaderboards can hide differences among tasks and categories.
Compliance is not graded on a curve
A student can pass a course with an 80% average because a strong final offsets a weak midterm. Nobody averages your compliance regimes.
An auditor examining your EU AI Act posture does not credit your excellent GDPR mapping against your missing risk classifications. A DORA examiner does not care that your SOC 2 evidence is immaculate. Each framework is scored on its own terms, against its own obligations, by its own reviewer. Strengths in one regime cannot pay down weaknesses in another. A gap in one regime is a gap in that regime, full stop.
A single aggregate benchmark answers a question no auditor asks.
Two benchmarks can share an average and hide opposite risks
Imagine two systems both score 80 percent. One gets 80 across every category. The other gets 100 on policy summaries, 100 on definitions, 90 on timelines, and 30 on breach obligations. The average is the same. The procurement decision should not be.
GRC benchmarks need framework, category, distribution, and coverage views. A tool that fails one obligation family can still look good in an aggregate. The buyer's job is to find the category that would embarrass them in front of an auditor before the tool does.
The research already said this
AI fairness evaluation research argues for benchmark suites that expose multiple measurements and trade-offs instead of reducing distinct concerns to one ranking.
NIST's draft guidance on automated benchmark evaluations says automated benchmarks are not well-suited to every use case. It tells evaluators to choose benchmarks that match the evaluation objective, inspect test items when possible, and report statistical analysis and qualified claims. A procurement decision needs more than the headline score.
What to demand from a benchmark
If a benchmark will inform a compliance decision, demand the breakdown:
- A score per framework, so EU AI Act performance is not hidden behind GDPR performance
- A score per category within each framework, so a strong average cannot mask a failed obligation family
- The distribution as well as the mean, so you can see where the model is uncertain
- Coverage stated explicitly, so a partial answer is not counted as complete
Sorena's benchmark reports results per category, per framework, and per session, with coverage tracked separately from the average.
A vendor who quotes only one number may not be measuring the categories or may not be showing the ones that fail.
The honest read of any score
When you see an aggregate benchmark score, ask which categories built the average and how far apart they are.
A tight band, where every framework lands near the mean, supports the aggregate. A wide spread means a few strong categories may be carrying failing ones.
Both distributions can produce the same headline score. The per-category breakdown tells you whether the model covers the framework you need. Insist on it.
Frequently asked questions
Is an average benchmark score ever useful?+
Yes, as a summary of the tested set. It still cannot show where the model fails. For a decision that depends on a specific framework or category, inspect that breakdown, its coverage, and the spread of results.
Why is averaging especially wrong for compliance?+
Because compliance is not graded on a curve. Auditors and regulators assess each framework on its own obligations. Excellent [GDPR](/artifacts/eu/general-data-protection-regulation) work does not offset missing [EU AI Act](/artifacts/eu/artificial-intelligence-act) controls. A single average assumes strengths can pay down weaknesses, which is precisely the assumption compliance does not allow. A gap in one regime stays a gap in that regime.
What should a trustworthy benchmark show instead of one number?+
A score per framework and per category, the distribution rather than just the mean, and coverage tracked as its own explicit metric. The goal is to make weak categories visible, not to compress them into a comfortable headline. A benchmark that only offers one aggregate figure is hiding the results that matter most.
Sources
- Wang et al., Benchmark suites instead of leaderboards for evaluating AI fairness (Patterns, 2024)https://www.sciencedirect.com/science/article/pii/S2666389924002393?ref=sorena.io
- The problem of aggregation, ML Benchmarkshttps://mlbenchmarks.org/12-problem-aggregation.html?ref=sorena.io
- NIST AI 800-2 (Initial Public Draft), Assessing and Managing AI Riskshttps://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf?ref=sorena.io


