On this page
Certified numerical computation reproducibility workbench
Edition 9 October 2026 v2. Retain a specialist evidence integration at low commercial confidence. There is a concrete artifact/claim review problem, but existing reproducibility tools are strong baselines. No paid demand, new algorithm advantage or independent source-proof acceptance is established.
Research finding
Family 266's stronger manuscript claims exactly three MUBs in complex dimension six under a documented binary64 computation. Its proof depends on geometric/arithmetic arguments and a complete covering/fallback execution. The exact Fourier companion supplies pair/mixed moment certificates and a no-seven-family argument; a separate completion theorem yields at most five. The selected MUBSix interface states Fourier vanishing and at most five. The separate cube-fiber interface supplies a supporting cancellation lemma. These are different conclusions.
The main PDF names eleven package paths that are absent from its directory in the pinned tracked tree and local checkout. This prevents reproducing that native pipeline here; it is not proof of absence everywhere or a mathematical counterexample. The exact companion's verifier is present and its SHA matches the paper and wrapper.
Problem and buyer
Research-software teams and laboratories preparing computer-assisted results need reviewers to determine which data, runtime assumptions and semantic arguments support each conclusion. A supervisor exit and a hash list alone can obscure incomplete coverage or a weaker formal theorem. This is an observed documentary gap; buyer frequency, willingness to pay and commercial materiality are unmeasured.
What the finding could enable
Extend the shared evidence platform with artifact availability, build-bound arithmetic obligations, stage/fallback coverage and claim-to-certificate mapping. The user receives a reviewable obligation graph and explicit unresolved items. A workflow can retain 163 initial graph-cap failures as pending until their fallback descendants are all accounted for, rather than treating caps as mathematical exclusions or all intermediate nonzero exits as final failure.
Technical and commercial limits
The native contract requires separate correctly rounded binary64 operations, nearest rounding including pre-main initialization, no contraction/fast-math reassociation, no FTZ/DAZ and masked exceptions. Two cover power expressions require the analyzed multiplication lowering in the actual build. Probe samples, compiler flags, containers and missing dynamic symbols cannot alone certify the whole arithmetic model. Rebuilt native hashes need not match historical hashes.
The reported 4 October run has 100 invocations and 4,470.33 seconds supervisor wall time. Neither it nor the earlier September execution was reproduced here. The exact verifier was not run, its moment identities were not rederived, and Lean was not executed. Independent exact positivity of six stored matrices is only one obligation. Specialized review effort may overwhelm revenue.
Minimal architecture
Pinned source/package inventory -> claim-specific obligation manifest -> source/binary/runtime identities -> terminal-state and record-provenance review -> domain certificate checks -> human-readable report with unresolved items. Use an established execution/package tool where appropriate. Bind any future replay to fresh output paths and preserve all failures. The current artifact is a static obligation review and independent matrix-data experiment, not a general trusted pipeline runner.
Existing alternatives and differentiation
BenchExec, reviewed 9 October 2026, already handles Linux benchmarking, subprocess resource measurement/limits and result reporting. ReproZip, reviewed the same day, packages research programs and dependencies. The upstream exact companion already contains a checksum wrapper and verifier. None was benchmarked here. Proposed differentiation must be useful domain-specific claim/coverage review beyond packaging; a generic wrapper is insufficient.
Monetization hypothesis
Retain AUD 5,000-15,000 for a bounded specialist audit as an unvalidated hypothesis within the shared platform. At an assumed AUD 8,000 and 30 hours at AUD 180/hour, AUD 2,600 remains before other costs; 45 hours cost AUD 8,100 in labor alone. Artifact recovery, platform-specific disassembly and mathematical review could consume more time. No outreach or paid pilots occurred; do not add this as independent recurring revenue.
Validation experiment
The source-data review adds 35 passing controls. Six literal Gram blocks have sizes 5/21/11/27/22/12. Independent exact LDL reconstructions verify 98 positive pivots; reversed-order pivots also satisfy the paper's greater-than-40,000 condition. Mutated and singular/indefinite controls are rejected. Source hash, declared count arithmetic and the displayed negative scalar are checked. No full moment identity or MUB bound is established by these tests.
A commercial trial would compare the report with ordinary package/CI review on permissioned projects, measuring actionable gaps, false alarms and total analyst time. Tests authored here cannot establish that value.
Conditions to reject or defer
Defer any verified exactly-three badge while the native package, terminal provenance, arithmetic conditions and semantic arguments remain unverified. Defer the exact companion's accepted-certificate label until the full calculation and model bridge are checked. Reject a standalone business if existing CI/reproducibility tooling and internal review already meet the need or specialist delivery exceeds fees.
Next concrete action
Preserve the missing-path inventory and request an authoritative package only through authorized access; do not invent absent binary inputs. For independent future work, measure the exact companion's full verifier in an appropriately bounded environment after complete code review, then verify its mathematical mapping. Continue the facility-review product, addressing the public planning workflow's weighted/rectangular input gap before polishing unsupported models into the interface.
Source evidence and reading extent
Review and source hashes, arithmetic/stage obligations. Main PDF pages 1-4/49-56; exact companion pages 1-3/16-25. Six pages visually inspected. Complete scope/Comparator configurations/wrapper/entry files and selected verifier functions read; literal integers parsed as data. This is selected-section review, not full proof acceptance.