Evidence Assurance Benchmark

A score we designed to be hard on ourselves.

Most vendor scores exist to look good. This one exists to be checked — a frozen, versioned benchmark that measures how the installed SBOMFlow product handles evidence: whether it proves what it claims, refuses what it cannot know, stays deterministic, and fails loudly instead of reassuringly. It was built so that a flattering-but-invalid result is structurally harder to produce than an honest bad one.

A benchmark that cannot hurt its author is not a benchmark, so the rules below were fixed before any result existed and cannot be tuned to one afterwards.

This space is reserved for a number we can prove.

Current public result

NO QUALIFIED CANONICAL RECEIPT — NO PUBLIC SCORE

Benchmark
Benchmark v1.0.0
Scoring policy
policy v1.0.6 — frozen before any measurement
Execution coverage
Qualification
NOT REACHED
Score
WITHHELD

A number appears on this page only when a run reaches QUALIFIED: every mandatory case executed, every integrity control positive, the tested artifact authenticated by digest, and the receipt independently re-verified. No run has reached that bar yet — and under the frozen policy, nothing weaker may carry a number.

Which product version does this cover? None yet. There is no qualified receipt for v0.5.0, none for 0.4, and none for any other version. The benchmark version and the scoring policy are shown instead, because both were frozen before any measurement existed. Release verification — how the v0.5.0 build was checked before it was handed to anyone — is a separate instrument with a separate purpose, recorded in the trust centre; its results are never a benchmark score.

The obvious objection

Why trust a benchmark written by the vendor it measures?

You shouldn't — not on our word. So the design assumes the party being measured is actively trying to obtain a flattering score, and closes the ways it could:

The rules freeze before the game starts

The dimensions, their weights, the critical failure classes and the qualification bar were set and versioned before any case ran. Changing any of them after seeing a result invalidates every receipt scored under that version.

Every check must prove it can fail

Each case carries a planted fault the check must flag and a healthy input it must pass. A check that cannot be made to go red for its intended reason scores nothing — only then does its pass count.

Public core, held-out cases

The public cases are fully published — fixtures, expected behaviour, scoring rules — so anyone can rerun them. A second set of cases is held outside the product's reach, so the product cannot simply be tuned to a corpus everyone can read.

Silence never qualifies

A separate verifier recomputes every figure in the receipt from the case-level records; any disagreement refuses the run. Attempted, executed, scored and skipped are four separate counts, published together — a skipped case is never a passed case, and “no violation was recorded” never upgrades a run on its own.

How one run works

From installed product to verifiable receipt.

The benchmark never measures our source tree — it measures the installed artifact, exactly as a customer would run it, offline, in isolation.

01

Authenticate the subject, freeze the inputs

The wheel under test is digest-verified and installed into a clean environment; every fixture is held to a recorded digest before a case may run. If the executable is not provably the artifact being claimed, or an input is not the declared input, the case ends as unmeasurable — never as a quiet substitution.

02

Execute in isolation

Each case runs the real CLI in its own working directory, offline, with the held-out cases outside the product's reach. Attempted network access — attempted, not merely successful — is itself a recorded failure.

03

Judge, then score under the frozen policy

Each check reads what the product actually wrote and answers one question, beside the denominators that make the answer mean something: how many items it examined, and whether its planted controls fired. Points earned over points eligible, per dimension, weighted as frozen — and a single critical violation (invented evidence, a silent loss, a leaked path, a nondeterministic canonical artifact) disqualifies the run regardless of everything else.

04

The number unlocks at QUALIFIED

Every count, verdict, digest and limitation lands in one canonical receipt, and a separate verifier recomputes it from scratch. Only a receipt that survives that, on one authenticated artifact, in one run, may ever put a number on this page. Until then the seat above stays empty.

What gets measured

Seven dimensions, weighted before any result existed.

One number would hide too much, so the score is a weighted composite of seven separately-reported dimensions. The weights below are read from the frozen, published scoring policy — this page cannot state different ones without failing its own build.

Evidence correctness & provenance

25

Does every recorded fact — a digest, a reference, a claim of origin — match the actual bytes?

Vulnerability intelligence

20

Are findings matched, attributed and suppressed only by the rules — with human decisions honoured and unattributed decisions refused, loudly?

Standards & interoperability

15

Do emitted documents hold to the formats they claim — CycloneDX, SPDX, OpenVEX — under an external validator, not our own opinion of our output?

Failure honesty & adversarial resilience

15

Given malformed, hostile or unscoped input, does the product refuse visibly — or absorb it into a reassuring answer? A crash is not a refusal; silence is not safety.

Determinism & reproducibility

10

Same inputs, same bytes — across repeated runs and across checkout locations.

Scale & resource discipline

10

Do declared memory, time and output budgets actually bind, and does an over-budget input degrade into a named warning instead of a crash or a silent partial result?

Installed-product usability

5

The mechanical floor: clean install, working entry point, honest exit codes, recoverable first run. Deliberately the smallest weight: necessary, not sufficient.

Read this before any number

What this score is not.

It is not a security percentage. The benchmark tests how the product handles evidence — never what that evidence says about anyone's security, and never whether a release is free of vulnerabilities.

It is not a regulatory result. No case outcome is evidence of conformity with the CRA or any other regulation, and it is not an endorsement by any standards body.

It does not generalise. A result binds one artifact, one commit, one machine, one benchmark version. It says nothing about other environments, other versions, or the parts of the product the declared cases do not reach — which is why execution coverage is published beside it.

It is not comparable across versions by default. Two results compare only under an identical scoring policy and a compatible corpus; anything else is labelled NOT COMPARABLE rather than plotted as progress.

It is not the release-verification result. How a release build was checked before it was handed to anyone is a different instrument answering a different question, and its outcome is never convertible into a score on this page. Two instruments produce two results, never one.

Check us

The public core is yours to run.

The benchmark contract, frozen scoring policy, public-core cases, runner and independent verifier are maintained alongside the product — not inside the installed package — with a methodology that records every anti-gaming control and every corpus amendment, and are available as a reproduction pack on request. When a qualified receipt exists, it will be published here with the exact commands to re-verify it. Until then, the only honest number is no number — and the product the benchmark measures is the one you can read about and, as an approved tester, run on a real release.