WORKBENCH
KALIBRA
TRUSTWORTHY VISUAL INSPECTION
OFFLINE MODE
model.onnx  0437ae28…741a copy
STATION 00 / OVERVIEW

Kalibra inspects — and decides whether its own inspection can be trusted.

An offline, reproducible visual-inspection runtime for industrial quality control. Every result on this page is backed by a governed, inspectable artifact — and everything not yet demonstrated is labeled as such.

Governed model artifact
sha256:0437ae28…741a copy
RUNTIME EQUIVALENCE: VERIFIED · REPLAY: PASSED
6,492 samples · max deviation 7.1e-15 · byte-identical second run
Single-seed, VisA-proxy evidence. Calibrated confidence — not yet demonstrated. This page never claims trust it has not proven.
Inspection · The runtime, followed end to end
A real governed decision

The runtime reaches a defect judgement and locates it. It does not decide whether that judgement can be trusted — the raw anomaly measure and calibrated confidence are never blurred.

pcb4 / anomaly / 049.JPG Recorded governed case
split test category pcb4 sha256 d7873fe6…a91d2e
Governed VisA-proxy PCB inspection input
Governed input pcb4 / anomaly / 049.JPG
Semi-transparent PaDiM raw anomaly intensity over the governed PCB input
PaDiM anomaly overlay Raw anomaly intensity over governed input — not calibrated confidence lowhigh raw anomaly
PaDiM raw anomaly map for the governed PCB case
PaDiM raw anomaly map map sha256 41c9f224…9d83e2
Recorded localization region overlaid on the governed PCB input
Recorded localization LOCALIZATION · x 0.250–0.375, y 0.375–0.500
Inspection result
DEFECT

Judgement only — “is this defective, and where?”

Raw anomaly measure — not calibrated confidence
6.276
010+ · model raw scale

A standalone measure of how unusual the input is — not a probability of being correct, and not calibrated confidence.

Precomputed from a governed offline run on the VisA proxy dataset. Kalibra does not run inference in your browser — this single case is illustrative evidence, not proof of generalized performance.
Walk the decision · governed identity
Model identitykalibra-padim-onnx-export-v1
Model artifactsha256:0437ae28…741a
Input content hashd7873fe6…a91d2e
Anomaly-map hash41c9f224…9d83e2
Projection manifestc60a3bd4…3bb25c
Feature contract…rgb64-bilinear-float64-patch8-v1
Toolchain pinonnxruntime 1.19.2 · numpy 2.0.2 · python 3.9.6
Determinismsingle-thread · ORT_DISABLE_ALL · CPU EP
Trust qualification — not yet demonstrated

This is where the raw measure would become calibrated confidence and an accept / review / reject / abstain outcome. Kalibra has not built or evidenced that layer, so no gauge is shown here. The absence is the honest state — not an error.

Calibrated confidence · absent Outcome routing · absent Drift · absent
What the runtime does carry
Learned PaDiM signal, end to endreal
Governed model artifacthash-anchored
Deterministic replaybyte-identical
Placeholder on canonical pathretired · false
Evidence · A chain to walk, a claim to verify
From raw dataset to a governed runtime

Evidence here means provenance and reproducibility, not scores. Every link is a claim in one sentence, a proof handle, and an invitation to re-check it against the repository yourself.

Equivalence · The runtime carries exactly the validated signal
Proven to machine precision

A serious runtime engineering result: across 6,492 samples the runtime's outputs match the offline-validated signal at the floor of float64 precision — far below the verification tolerance.

Maximum runtime deviation vs. tolerance
SAMPLES COMPARED
6,492
MAX ABS DEVIATION
7.105e-15
1e-18
1e-15
1e-12
1e-9
MAX DEVIATION
TOLERANCE 1e-12

The maximum deviation sits over 140× below the 1e-12 tolerance — at the noise floor of double-precision arithmetic. In plain terms: the runtime carries exactly the validated offline signal.

tolerance atol 1e-12 rtol 1e-12 localization exact · 0.0
Deterministic replay · run it twice, byte for byte

Same input, two independent runs. Every governed comparison in runtime_replay.json is identical.

Artifact identity✓ true
Predictions✓ true
Localization✓ true
Raw anomaly measures✓ true
Run hash✓ true
Session config✓ true
Output digest✓ true
status: passed 7 / 7 comparisons
Architecture · A decision flowing through a system
Five domains — complete where real, unbuilt where honest

A decision is a chain, not a verdict. This flow is complete from Inspection → Evidence → Evaluation, and deliberately unbuilt at Trust → Review. The gap is the honest story, drawn to scale.

Real
Inspection Engine
What the system sees. Carries the governed PaDiM signal end to end and emits a raw anomaly measure — explicitly not confidence.
Not yet
Trust Qualification
How far to trust it. Where a raw score would become calibrated confidence and accept / review / reject / abstain. Designed as a contract — not yet evidenced.
Not yet
Human Review
Where uncertainty goes. An architectural seam only — the interactive review loop is not yet demonstrated.
Real
Evidence Engine
What can be inspected. The governed artifacts and SHA-256 hashes the whole page stands on.
Real
Evaluation Engine
What the science says. Governed C-6 metrics, reported honestly — weak numbers included.
Inspection Trust Qualification Human Review Evidence Evaluation
Trust · Inspectable, not authoritative
Why trust this result?

Not a gauge. A plain statement of what is demonstrated, what is not, and how a skeptic can verify every claim without reading the code.

What is demonstrated
  • A real learned PaDiM signal carried end to end by the runtime.
  • Runtime equivalence to machine precision across 6,492 samples.
  • Byte-identical deterministic replay across all seven comparisons.
  • A hash-anchored evidence chain from dataset archive to runtime.
What is not
  • No calibrated confidence — the raw measure is not a probability of being correct.
  • No accept / review / reject routing and no abstention.
  • No drift assessment and no interactive human-review loop.
  • Single seed, VisA-proxy domain — not a production or domain-of-record claim.
How to verify
  • Copy any SHA-256 on this page and re-check it against the repository.
  • Regenerate the equivalence and replay records from fixed inputs.
  • Read the governed evidence documents linked at Station 08.
  • Every absence above is stated — nothing is hidden in a footnote.
Boundaries · Honesty read as rigor
The science, reported honestly

Governed C-6 metrics, each carrying its qualifier inline. The weak numbers are shown, not hidden — evidence that the author reads their own results honestly.

Image AUROC
0.757826
Detection quality at the image level.
Pixel AUROC
0.865196
Localization quality at the pixel level.
AUPRO
0.555765
Per-region overlap — the harder, more honest localization measure.
Every metric above is single-seed, VisA-proxy only The qualifier travels with the number, always. Diagnostic weakness is named too — per-class precision falls as low as 0.209 (high false positives). It is shown, not buried.
Not yet demonstrated — one consistent, calm convention
Calibrated confidence
Accept / review / reject routing
Abstention on low confidence
Drift assessment
Interactive human-review loop
Multi-seed variance
Multiple model families
Real-time / on-device inference
Production / deployment

A clearly drawn limit is a sign of understanding, not a gap. The trust-qualification layer is the thesis of the whole system — and Kalibra says, in the same calm voice it uses for its proofs, that it is not yet demonstrated.

Timeline · Governance as a build history
How it was built, and proven

Process is only interesting once the output is believed. The phase lineage turns the repository's own governance discipline into an inspectable artifact.

P1
Engineering substrate
Governed foundation, five-domain architecture, reproducible offline boundary.
Offline, batch, locally reproducible system boundary defined.
Five engineering domains established as architectural boundaries.
SHA-256 governance introduced across every artifact.
P2
Offline science
Governed VisA acquisition, PaDiM baseline, C-6 scientific evaluation.
Governed VisA proxy acquisition (visa-acq-v1) — archive + split hashed.
PaDiM baseline fit (visa-padim-baseline-fit-v1) with training replay.
C-6 evaluation: Image AUROC 0.757826 · Pixel AUROC 0.865196 · AUPRO 0.555765 (single-seed, VisA-proxy only).
P3
Runtime integration
ONNX export, export- & runtime-equivalence, replay, placeholder retirement.
Governed ONNX export (opset 18, IR 10) + export-equivalence vs PyTorch.
Runtime equivalence: 6,492 samples, max deviation 7.105e-15 (tolerance 1e-12).
Deterministic replay status: passed7/7 comparisons byte-identical.
Canonical placeholder retired — placeholder_used_on_canonical_runtime_path: false.
Next Trust qualification — calibration, routing, abstention, drift, human review. Designed, not yet evidenced.
Verify · The exit is a re-check, not a signup
Reproduce it yourself

The page earns the click to the repository by first proving the work is worth reading. Every claim regenerates from a fixed starting point.

Level 1 · Public clean-clone verification

Works from a normal public clone using tracked repository contents only.

verify — public clone
python3 -m pytest -q
python3 scripts/verify_public_clone.py
python3 scripts/build_portfolio_evidence_bundle.py --check
Level 2 · Governed-data verification

Requires separately acquired governed VisA data and is expected to fail closed when that data is absent.

verify — governed data
python3 scripts/verify_padim_runtime_equivalence.py verify
python3 scripts/verify_placeholder_retirement.py verify
python3 -m pytest -q -m governed_data

Level 3 · Full scientific reproduction is a separate workflow after governed acquisition; follow the repository methodology.

Kalibra repository
Static site lives next to the source · GitHub Pages · evidence review HEAD d8bba98
Commit against which the displayed evidence bundle was reviewed; not the current repository HEAD.
Where the evidence lives
Scientific evaluation (C-6)docs/evidence/…EVALUATION
Runtime equivalencedocs/evidence/…EQUIVALENCE
Runtime provider integrationdocs/evidence/…INTEGRATION
Placeholder retirementdocs/evidence/…RETIREMENT
Runtime artifactsartifacts/runtime/*.json
Regenerable from fixed inputs · no live inference, ever.
Static portfolio surface · no CDN, no external libraries, no live inference. Fonts reference the Kalibra identity (Space Grotesk / IBM Plex) with system fallbacks. All values shown are governed artifact values or explicit “not yet demonstrated” states — no fabricated metrics.