How to Benchmark AI for Compliance

Every AI vendor cites a benchmark number. Few of those scores show whether the tool can finish a compliance task an auditor would accept. Test the work you plan to buy the tool for.

Sorena AI TeamResearch and Benchmarks4 min read

The score is not the job

Vendors often present a benchmark score as evidence of capability. The score only shows performance on the task and metric that the benchmark was built to measure.

Many public AI benchmarks compare general-purpose models on general-purpose tasks. A compliance buyer needs to know whether a tool can complete a real obligation correctly, cover every required element, and support its claims with sources.

Stanford HAI reviewed more than 50 benchmarks and found that many had saturated around 80 to 90 percent accuracy, with models above human baselines on some tests. It also warned that models can perform well on those benchmarks and still give incorrect answers in use.

Why generic scores do not establish domain competence

MMLU is a bank of multiple-choice questions across many subjects. A model selects from predefined answers and receives credit for the correct choice.

That format does not test the full process of mapping ISO controls to evidence or analysing whether a regulation applies to an entity. Those tasks are open-ended and require the system to produce and support every relevant part of the answer.

A model can perform well on MMLU and still miss items in a control mapping because the tasks test different capabilities. Treat a generic score as background evidence, then test open-ended production on the target workflow.

Fluency is not completeness

A fluent answer may still be incomplete. General-purpose models can produce authoritative prose without covering every part of the task.

A response that addresses six of twelve obligations leaves half of the required set uncovered. The prose may still look finished, so reviewers need a checklist that exposes what is missing.

Measure completeness against a fixed set of expert-defined requirements. Score how much of the required work is present and whether each claim traces to a source. Sorena's benchmark work on real GRC tasks uses coverage against an auditor checklist.

Use a buyer-ready benchmark rubric

Before trusting any AI benchmark, ask for the rubric. The minimum credible package is the task prompt, expected answer elements, source set, scoring criteria, reviewer qualifications, baseline model version, run date, coverage denominator, and examples of failed answers.

If those details are missing, the score is a marketing asset, not evidence. A compliance benchmark should make weak spots inspectable. You should be able to see whether the system missed obligations, cited weak sources, guessed around uncertainty, or produced a correct answer for the wrong reason.

Real tasks, scored by real experts

Use tasks drawn from the actual job: a privacy audit, framework crosswalk, applicability analysis, or clause-by-clause comparison. These expose omissions that a multiple-choice test cannot.

Have qualified reviewers score the output against a disclosed rubric. Model grading can be useful, but an independent domain expert should decide whether the work is complete and defensible. NIST SP 1270 treats AI bias as a socio-technical issue and emphasizes context of use, human factors, stakeholder involvement, and task-specific evaluation. A vendor should be able to explain who scored the benchmark and how.

How benchmarks get gamed

Three problems can weaken a benchmark:

  • Contamination. If test questions appeared in training data, the model may recall answers rather than solve them.
  • Format overfitting. A model tuned for multiple choice may perform differently on open-ended production work.
  • Saturation. When scores cluster near the top, small differences may say little about practical performance. Stanford HAI documented ImageNet moving from 91 percent to 91.1 percent in a year.

Use fresh tasks, an open-ended format that mirrors the job, and independent expert scoring.

What a real compliance benchmark looks like

A useful compliance benchmark includes:

  • Tasks drawn from real GRC work.
  • An open-ended format that requires the full answer.
  • A fixed checklist of expert-defined requirements.
  • Independent domain experts as reviewers, ideally in more than one pass.
  • Completeness, source grounding, and factual errors reported separately.
  • Tasks fresh enough to reduce contamination risk.

Sorena AI's benchmarks use real compliance tasks and checklist-based scoring. Stanford HAI likewise recommends broader evaluations such as HELM, which measure accuracy alongside robustness and other factors, because a single saturated accuracy score gives an incomplete picture.

How to read a vendor's benchmark number

Ask what tasks were used, who scored the output, whether completeness was measured against a checklist, whether the test data could have appeared in training, and whether you can reproduce the result on your own work.

Then run the tool on a representative task, have a qualified reviewer check the output, and count the missing requirements. Your own evaluation is the most relevant evidence for your workflow.

Frequently asked questions

Are benchmarks like MMLU useless?+

No. They are useful for what they were designed for: comparing general-purpose models on broad, multiple-choice knowledge. They just do not predict domain competence. A high MMLU score tells you a model is broadly capable, not that it can produce complete, defensible compliance work on open-ended tasks.

Why do experts have to score a compliance benchmark instead of an automated metric?+

Because completeness and defensibility are judgment calls that depend on context. An automated metric can check whether text matches a reference, but it cannot reliably decide whether an obligation was fully addressed or whether a claim is traceable to a valid source. NIST SP 1270 treats AI evaluation and bias management as socio-technical work involving context of use, human factors, stakeholder involvement, and task-specific evaluation, not just a single automated score. For compliance work, the expert is the practical ground truth.

How do I know a vendor's benchmark was not gamed?+

Ask whether the tasks were fresh and open-ended, who scored the output, and whether you can reproduce the result on your own work. Also inspect the rubric and coverage denominator. Vague methodology makes the score difficult to assess.

Sources

Share

See Sorena do the work

Book a demo and watch one real compliance workflow go from question to audit-ready output.