Adversarial test scorecard

Adversarial lab coverage demonstrates engineering robustness only. It is not a security assurance, not a certification, and not evidence of legal or CRA conformity.

This page describes SBOMFlow 0.4.0. The current build is 0.5.0. The run recorded here was executed against an earlier revision of the repository, at the version and revision recorded under The run below, and it has not been re-run since. No adversarial-laboratory run has been recorded for 0.5.0, and none is simulated here: a page that showed numbers for a run that never happened would be a fabricated measurement, so this page keeps the run it actually has. What it establishes is what that earlier run found, at that revision. What it does not establish is the behaviour of the current build, and it is not evidence about 0.5.0 in either direction — neither that the findings still hold nor that they were fixed. How the 0.5.0 build itself was checked before it was handed to anyone is a separate record, in the trust centre.

Every number on this page is generated from a recorded lab run. Nothing is typed in by hand, and a drift guard fails the build if this page and the recorded run disagree.

The run

As of
2026-07-03T12:00:00Z (pinned reference date, not the execution time — the snapshot carries none, so the artifact stays reproducible; the version and revision below are the freshness signals)
Tier
standard
Result
pass
Offline
True
Seed
20260703
SBOMFlow version
0.4.0
Repository revision
a514f59ca353bd775be04b3d2e68d8988b47922a

Cases by area

Cases recorded by lab area in this run.
AreaCasesPassedFailed
customer-isolation110
determinism110
identity-accuracy880
launch-readiness330
license-accuracy110
release-gate10100
resilience10100
security-redaction10100
usability-contracts880
vulnerability-accuracy990
Total61610

Determinism

This recorded run used ONE repetition, so it did not measure whether repeated runs produce byte-identical artifacts. The digest below covers the artifacts of that single run. The laboratory supports repeated-run comparison; this snapshot simply does not contain one.

Artifacts digested
44
Cross-repetition byte identity
not measured — this run used a single repetition
Basis
audit(--as-of pinned, fixed probe path per invocation)

Mutation self-checks

Deliberate faults injected to prove the lab can actually fail. A suite that never fails when the product breaks is not evidence.

Executed
True
Controls detected
4

Budget

Observed duration
65.1s
Tier budget
900s
Within budget
True

Known gaps

What this run did not establish. Recorded because a scorecard that only lists wins is marketing, not evidence.

Capability coverage

56 of 56 capabilities recorded in the lab's coverage matrix are implemented and exercised by at least one named case. The matrix records the capabilities this lab tracks; it is not a claim that it enumerates everything a given product needs.

This page describes engineering testing only. It is not an audit, not a certification, and not a statement about any specific product's regulatory position. How we run this laboratory is documented in the trust centre; the scoring discipline for the engine itself lives on the Evidence Assurance Benchmark.