AI Evidence

Know what intelligence shipped.

Ordinary SBOMs describe software components. AI-enabled products also depend on models, datasets, prompts, services, agents, and evaluations. SBOMFlow is connecting that evidence to the same release record, review process, gates, and history as the rest of the product.

In development · design-partner programme First foundations shipped in 0.4

Straight status, up front: this page describes a capability we are building in the open with design partners — not a finished product surface. What already runs today is labelled available; everything else is labelled exactly as far along as it really is.

The questions this exists to answer

When someone asks what your product's AI actually is.

Which AI-enabled features shipped in this release?

Not in the design docs — in the release. Tied to the same release identity as every other component.

Which models or external AI services were used?

Model files observed in the build, and declared external services — recorded separately, so a declaration never impersonates an observation.

Which datasets, prompts, agents, and tools were declared or observed?

Where they are genuine release artifacts, they belong in the record with provenance.

What changed since the last approved release?

A new model file, a swapped external service, a changed prompt bundle — release-to-release drift, extended to AI evidence.

Which provenance, licence, evaluation, or documentation gaps remain?

Gaps stated plainly, as engineering facts — the same discipline the rest of the record already has.

Who reviewed and approved the material AI change?

The same human review queue, approvals, and explainable gates — extended, not duplicated.

What can be shared without leaking sensitive IP?

A redacted, verifiable sharing output for customers and auditors — the pattern our sharing packs already follow.

What runs today

The honest starting point.

Available today

Model files are never ignored

The standard scanner already recognises model-artifact families — SafeTensors, GGUF, ONNX, TFLite, Core ML, OpenVINO, TensorRT, and pickle-class checkpoints — in any scanned release: detected, hashed, and surfaced as recognised evidence with a warning, deliberately never parsed or loaded.

Shipped in 0.4 · foundations

Release-grade model identification

The 0.4 release line adds the AI evidence foundation: model artifacts identified from bytes — documented magic numbers and structural invariants, never execution — with identification strength recorded per artifact, so a claim resting on a filename is never presented as a claim resting on bytes. Formats that can only be read by executing code are recognised and refused. Operator-facing commands are the next step, in engineering now.

Available today

External AI-assisted scans as labelled observations

Sealed output from an external AI-assisted security scanner can be imported from local disk: every finding arrives as an external, AI-assisted observation that needs human review — never a SBOMFlow finding, never a gate input, and never re-scored. Partial coverage stays visible, and zero findings never means safe.

The evidence chain

What the AI release record covers — and where each piece stands.

Statuses below are enforced by the same no-overclaim discipline as the rest of the site. When a row changes state, this page changes with it — not before.

EvidenceWhat it recordsStatus
Model artifacts in the release Files observed in the build: format, hash, size, identification strengthRecognised and hashed today; byte-level identification shipped in 0.4 as foundations Foundations shipped
Model & external-service inventory Declared models, deployment aliases, and external AI services, reconciled against what was observed In development
Datasets & provenance Dataset identity, provenance, licences, transformations, declared limitations Planned
Prompts & prompt bundles Prompt artifacts where they are genuine release artifacts, hashed and versioned Planned
Agents, tools & MCP dependencies Agent workflows, tools, and server dependencies with declared permissions and authority Planned
Frameworks, runtimes & supporting software The conventional software around the models — already first-class composition evidence Available today
Evaluation & test evidence Evaluation results, thresholds, and human acceptance recorded as evidence, never as verdicts Planned
AI-specific release drift Model, service, and prompt changes between releases, on the existing drift engine In development
Human review & AI-aware gates The live review queue, approvals, and explainable gates extended with AI-aware policies In development
Standards-based AI/ML BOM output Linked software SBOM plus CycloneDX ML-BOM and/or SPDX 3.0 AI & Dataset profile output Planned
Redacted AI evidence sharing A verifiable, redaction-audited AI evidence pack, following today's passport pattern Planned

Standards named exactly: CycloneDX ML-BOM and the SPDX 3.0 AI & Dataset profiles are the output targets we build toward — SBOMFlow defines no proprietary standard, and no generated document is ever regulator-approved.

Boundaries, from day one

The rules this capability is being built under.

  • Nothing executes, imports, or deserialises a model — model files are inspected as bytes, and pickle-backed formats are detected and refused, never parsed.
  • Observed is not declared — what a human or supplier asserts stays separately attributed, so a conflict is visible instead of silently overwritten.
  • Identification confidence is recorded, never rounded up — a magic-number match and a filename match are different claims and stay different.
  • Offline — a model reference naming a registry is recorded as a reference, never fetched.
  • Zero findings is not zero risk — an absent model is an absent observation, and an unreadable file is neither absent nor clean.
  • AI evidence is evidence about the product — SBOMFlow's own evidence path stays deterministic, with no required AI calls.

Where this matters first

Practical products, practical questions.

This is release assurance for real connected products with AI inside — not speculation about artificial general intelligence.

Smart cameras & edge inference

On-device detection models that change between firmware releases — often without anyone recording which weights shipped.

Robotics & autonomous functions

Perception and control models whose provenance and evaluation evidence belong in the release record their safety case points at.

Industrial & building-control AI

Anomaly detection and optimisation models inside OT products, where operators ask exactly what runs on their network.

AI-enabled network appliances

Traffic classification and filtering models — security products whose own composition should be the best-documented of all.

Products on external model APIs

Where the "model" is a remote service: which service, which deployment alias, what data-handling context — declared, dated, and reviewed.

Any release a customer questions

When a buyer's security review asks "does this product use AI, and how do you know?" — an evidence-backed answer instead of a shrug.

Help decide what "AI release evidence" means in practice.

We're building this with a small number of teams that ship AI-enabled connected products. Design partners get the working capability as it lands — and shape the evidence model, the review boundaries, and the sharing format against their real releases.