Coverage Is the Number That Actually Matters.

A general assistant can write a polished answer about your GDPR obligations. If it names 27 of 30 expected items, three obligations are still missing. Measure coverage to expose those gaps.

Sorena AI TeamResearch and Benchmarks4 min read

We are scoring the wrong thing

People often judge an AI answer by whether it is clear, confident, and knowledgeable. That may be enough for an email, but compliance also requires completeness.

Fluency measures how good an answer sounds. Accuracy measures whether its claims are true. Neither shows whether the answer left something out.

An answer can be well written and accurate about every obligation it names while still missing a third of the applicable obligations. Coverage catches that failure.

Coverage is recall, and recall is the point

The metric has a name outside compliance. In machine learning, Google defines recall as the share of actual positives correctly classified: TP divided by TP plus FN. In information retrieval, NIST TREC evaluation uses recall and precision to ask how much of the relevant set the system retrieved.

Precision measures how many returned items were correct. It does not penalize an item the system never mentioned. Recall counts those misses.

Coverage applies recall to obligations. Of everything the task required, what share did the system find, map, and address? This tells GRC reviewers whether gaps remain.

Sixty percent is not a passing grade

A 60 percent result leaves 40 percent of the required work uncovered.

An auditor does not erase the three applicable obligations you missed because you handled the other 27. One unaddressed data-transfer clause, skipped breach-notification timeline, or unmapped control can still produce a finding.

Partial answers are hard to spot because missing obligations do not announce themselves. The answer can look complete until someone checks it against the full requirement set.

Always inspect the denominator

Always inspect the denominator. If a task has 30 expected obligations and the system finds 27, the answer is missing 10 percent of the work. Those three misses are false negatives.

Benchmark buyers should ask for the expected obligation count, found obligation count, missed items, and source passages behind each answer. Accuracy on the found items does not show what the system skipped.

The miss costs more than the mistake

A system can include something that does not belong or leave out something that does. Machine learning calls these false positives and false negatives, and their costs differ.

A false positive wastes reviewer time on an obligation that does not apply. A false negative leaves an applicable obligation out entirely, sometimes until a regulator, customer, or auditor finds it.

A fluent answer may not signal what it omitted. Measure false negatives directly instead of inferring completeness from the prose.

What the standards actually ask for

The NIST AI Risk Management Framework lists valid and reliable as the first characteristic of trustworthy AI. It treats reliability as correct operation under expected conditions over time, including hard cases.

ISO/IEC 27001, NIST SP 800-53, GDPR, the EU AI Act, and the EU Data Act define requirements, controls, rights, duties, and conditions. Handling some applicable obligations well does not cover the ones a system missed.

The evaluation must match the standard. If the standard demands completeness, the metric must expose coverage.

How to read a benchmark once you know this

Check three parts of any coverage result.

First, inspect the denominator: the complete, expert-defined list of items the system should have found. A percentage against a vague target has little value.

Second, count the missed items as well as the incorrect ones. A benchmark that tallies only wrong statements measures precision without showing what the system skipped.

Third, trace each obligation to a source. Our benchmarks score coverage against auditor-defined checklists and require each point to be grounded in a primary source.

Measure the outcome you need

You get the behavior you measure. Fluency rewards answers that sound finished. Accuracy rewards true statements, even when the set is incomplete. Coverage forces the system to show what it found and missed.

Ask every tool how much of the required work it caught, not only whether its output reads well. Reviewers will check coverage.

Frequently asked questions

Is coverage the same as accuracy?+

No. Accuracy measures whether the statements a system makes are true. Coverage measures whether it addressed everything it should have. An answer can be accurate on the obligations it names and still miss others entirely. Accuracy grades what is present; coverage counts what is absent.

Why is coverage the same idea as recall?+

Recall measures the share of actual positives or relevant items a system returns. Coverage applies that directly to compliance: of all the obligations expected for the task, what share did the system find and address. Both metrics penalize the miss, which is exactly what a compliance review does.

Does high coverage mean the work is done without human review?+

No. Coverage tells you the system found and mapped the obligations you owe. People still make the judgment and risk decisions on top of that. Coverage removes the most dangerous failure, the silent gap, so reviewers can spend their time on decisions instead of hunting for what was left out.

Sources

Share

See Sorena do the work

Book a demo and watch one real compliance workflow go from question to audit-ready output.