On this page
- Complete utility directory
- Newly completed statistical integration
- New scheduling references
- Subset-sum reference and existing-engine evidence
- Literal packaging and source-derived count stage
- Exact polynomial and numerical-range evidence
- Facility-planning evidence
- Current documentary evidence
- Edit-distance source activation and existing exact tools
- Ordered-algebra formula testing
- Game algorithms and strategy-region evidence
- First local facility-review release
- Evidence-package review and first-product fit
Thirty-seven runnable research utilities
Fifty-eight saved component reports contain 21,853 passed checks. This edition adds a dated source-data study with 35 controls on family 266, including exact reconstruction of six supplied Gram matrices. The utility count remains 37; a study script is not an additional product. No complete MUB verifier, native pipeline or Lean kernel was executed.
Complete utility directory
| Utility | What it does | Material boundary |
|---|---|---|
| proof_preflight.py | Inspect selected formal configuration/module metadata | Static presence/trust settings, no proof execution |
| withdrawal_impact.py | Record explicit withdrawals and dependencies | Three notices/two explicit edges; not a full theorem graph |
| audit_embedding.py | Check finite embedding quantities | Sampled/finite floating calculation, no universal guarantee |
| gad_capacity.py | Evaluate scalar generalized-amplitude-damping expression | Numerical convention/optimization, no hardware or finite decoder certificate |
| audit_claim_contract.py | Lint documentary source/model obligations | Twenty-one curated contracts; supplied statuses and evidence contents unverified |
| periodic_interface_reference.py | Periodic interface formula reference | Specified finite geometry, no general optimizer |
| audit_binary_waveform.py | Exact finite correlation and sampled spectrum | No full channel or physical waveform certification |
| audit_ramanujan_graph.py | Rational contrast-space spectral acceptance | Supplied finite graph, not source constructor |
| cyclic_chain_reference.py | Tiny group cyclic-chain enumeration | At most sixteen elements, no effective spectrum witness |
| evidence_bundle.py | Join source hashes/static metadata/claim audit | Archives six exact inputs; no source/proof execution |
| contingency_reference.py | Exact bounded table count/rank/unrank/law | Dimension six/total128; conventional DP, not source sampler |
| audit_matching_certificate.py | Check matching and attaining cardinality bound | Exact supplied witness/bound, no extra allocation constraints |
| solve_and_audit_matching.py | Existing NetworkX proposal plus separate audit | Unit cardinality model, no source accelerated matcher |
| perfect_matching_reference.py | Tiny exact counts/ranks/edge fragility | At most24 vertices, no FPRAS |
| audit_switch_chain.py | Exact tiny specified-kernel TV | At most6 vertices/512 states/64 steps; complete-host model |
| audit_tree_thinness.py | Exact all-cut finite tree audit | At most16 vertices; not broad thin-tree construction |
| kserver_reference.py | Rational offline optimum and named-policy replay | At most8 points/4 servers/64 requests; no source online policy |
| transport_sharpness_reference.py | Exact three-atom rational geometry | Uniform-square example, no arbitrary transport solver |
| audit_fourier_aliasing.py | Sparse exact odd-power convolution/grid folding | Typed support/work/bit caps; no PDE solution certificate |
| chromatic_basis_reference.py | Tiny exact chromatic/e-basis reconstruction | At most6 vertices; not broad packet/witness theorem implementation |
| queryable_permutation_reference.py | Shared-switch point/inverse replay and tiny laws | Explicit bits/dyadic domain, no universal mixing constant |
| permutation_provider_reference.py | Decimal-wire seeded or locked local stored bits | Distinct randomness models; capped append-only JSON store |
| palindrome_minorant_reference.py | Exact tiny palindrome law/operator calibration | Two/four slots; no universal numerical P |
| audit_contingency_law.py | Complete finite law and fixed-event comparison | Explicit conditioning, at most2,000 tables, unknown on exhaustion |
| bounded_flow_reference.py | Signed-bound private-vertex flow reduction/counts | Tiny exact table engine; no unrestricted FPRAS/sampler |
| three_machine_reference.py | Unit-job schedule audit/exact tiny optimum/deadline evidence | At most18 jobs/50,000 states/2m transitions; exponential |
| solve_and_audit_three_machine.py | Existing CP-SAT proposal and finite independent evidence | Same strict unit model; solver status distinct from independent optimum |
| subset_sum_reference.py | Exact disjoint-half support/multiplicity/witness reference | Two–64 positive identified items; 50k sum records/2m generation/probe work, not source backend |
| subset_sum_schedule.py | Exact known guards and conditional decision-error/amplification algebra | Fixed main cutoffs unselected; source implementations and independence unvalidated |
| solve_and_audit_subset_sum.py | Existing integer CP-SAT proposal and original-ID witness evidence | Deliberate total/target cap 2^62-1; solver satisfaction status does not give uniqueness/count |
| superstring_packaging_reference.py | Emit indexed immutable byte-view artifacts and compare complete compressed bytes | At most128 records/16KiB literal bytes; exact DP at most 12 reduced strings; no source factor-two backend |
| superstring_counts_reference.py | Bounded paper forced-count recursion and balanced base graph | Reduced length 128/512 substrings/2m charges; no full layer/request/cycle construction or proof acceptance |
| numerical_range_polynomial_reference.py | Exact finite norm threshold, PSD support halfspaces and covered polynomial upper bounds | n≤8/m≤4/degree≤6; new constant two conditional, exact inputs and capped arithmetic only |
| metric_facility_reference.py | Strict metric, exact plan/LP-dual audit and bounded k-median enumeration | At most 48 locations, unweighted and uncapacitated; no source backend or physical-data bridge |
| noncommutative_hitting_reference.py | Exact ordered rational matrix witnesses via source-derived structured actions | Visible division-free tree; nonzero refutes a free identity, zero remains conditional/unknown; explicit bit/work/dimension caps |
| mean_payoff_label_reference.py | Bounded deterministic source recursion and separate threshold-region strategy checker | Ordinary signed-edge deterministic model; unknown on limits, completed labels source-conditional; no value-optimality or companion claim |
| facility_review_workbench.py | Stable-ID CSV import, exact model/plan review, optional OR-Tools and preserved HTML/JSON exports | Local strict/unweighted/uncapacitated model; no hosted uploads, automatic general duals or validated customer value |
All dated validators/publishers and archived historical copies are supporting tooling, not additional utilities. The preceding prototype edition preserves detailed earlier measurements and controls.
Newly completed statistical integration
586 live SciPy controls compare the exact reference with SciPy 1.18.1/NumPy 2.5.3 on 284 two-by-two margin pairs and 494 observed tables, plus five public/generated cases. A working fresh environment was created without modifying the default or first failed fresh 1.15.3 environment. R remains unexecuted. Floating agreement is finite and tolerance-bounded, not an exact library proof or design validation. Runtime and source hashes, public law report.
python3 research/tools/audit_contingency_law.py research/fixtures/contingency-scipy-fisher-threshold-2026-10-09-v1.json
The declared probability-order event is held fixed. Conditional-independence probability 5/143 versus uniform-table probability 3/7 crosses the declared 1/20 threshold because the law changes. Bounds use explicit renormalization of the original law; they do not automatically model real structural zeros. No large-support or source-backend performance was measured.
New scheduling references
3,346 finite controls compare all 1,099 topologically ordered DAG edge sets through five jobs with a separate backtracking slot-variable oracle. The reference validates explicit model fields, job identities and acyclicity; checks capacities/strict precedence; computes elementary lower/checked upper bounds; and completes ideal-state BFS when needed. Empty, oversized, cyclic and mismatched inputs are refused. Budget exhaustion grants no exact optimum. Independently sufficient deadline evidence can survive exhaustion.
python3 research/tools/three_machine_reference.py research/fixtures/three-machine-greedy-counterexample-2026-10-09-v1.json
python3 research/tools/three_machine_reference.py research/fixtures/three-machine-sync-bottleneck-2026-10-09-v1.json --state-budget 1
The generated nine-job case has a four-slot priority heuristic versus optimum three. The six-job case has elementary lower bound two versus optimum three. A valid schedule alone is not an optimality certificate; a valid schedule attaining the elementary lower bound can be one for the finite model. The source exponent 150020 and fixed-machine proof are not implemented by this exponential search.
40 live CP-SAT controls compare eight generated models with OR-Tools 9.15.6755, retaining separately checked schedules and tiny optima. Runs with one worker and a two-second limit already reach those optima, so no new solver advantage is shown. A FEASIBLE proposal attains the capacity bound in the independent-18 case. The exhausted independent six-job control does not inherit its solver's OPTIMAL status. Saved comparison, wheel installation report.
/tmp/openai-math-scheduling-20261009-FKmLBr/bin/python3 research/tools/solve_and_audit_three_machine.py research/fixtures/three-machine-greedy-counterexample-2026-10-09-v1.json
That interpreter is the preserved environment for this run; the saved report records packages/runtime. No existing installation was altered. Times are diagnostic single runs, not a benchmark or deployment guarantee. Current scheduling dossier.
Subset-sum reference and existing-engine evidence
2,294 controls compare 28,404 method/target queries across all 1,089 positive value words 1..3 through six items with a separate binary-choice count oracle. Both exact reference methods preserve original-index multiplicity. Additional controls check invalid models/IDs, 256-bit exact arithmetic, guards, budget states and conditional majority-error algebra.
python3 research/tools/subset_sum_reference.py research/fixtures/subset-sum-repeated-64-2026-10-09-v1.json
python3 research/tools/subset_sum_reference.py research/fixtures/subset-sum-repeated-64-2026-10-09-v1.json --method occurrence_MITM
The 64-one target-32 generated case uses 98 peak sum records under compressed support and retains C(64,32)=1,832,624,140,942,590,534 distinct indexed solutions. The occurrence method exhausts 50,000 records and returns unknown. These record/work caps omit runtime objects, input and native sorting scratch/comparisons, so they are not byte, time or word-RAM bounds. The duplicate-value control has two valid selections; repeated IDs are refused.
150 live solver controls cover 1,080 tiny target queries and ten generated cases. Existing CP-SAT already proposes exactly checked selections for repeated64 and distinct-powers32. A directly checked positive witness can suffice when count/reference search is unknown. Satisfaction OPTIMAL is no uniqueness/count certificate; solver INFEASIBLE remains a solver report unless independent negative evidence exists. Larger Python-exact totals are refused by the conservative integer adapter without rounding. Saved comparison.
/tmp/openai-math-scheduling-20261009-FKmLBr/bin/python3 research/tools/solve_and_audit_subset_sum.py research/fixtures/subset-sum-duplicate-identities-2026-10-09-v1.json
The preserved isolated OR-Tools 9.15.6755 environment was reused without installation changes. These are single-run generated controls, not real accounting data or speed comparisons. Updated dossier.
The source-parameter plans expose known 64-bit main guards at 600,000 and 36 billion items, with further fixed cutoffs unevaluated. At n64 and conditional all-call error budget 1/1,000,000, 261 odd-majority repetitions are assigned per two-sided adaptive decision query. This requires fresh independent genuine base-error calls, not implemented source programs. A verified positive witness and probabilistic NO remain distinct. Fast time and companion low-space claims cannot be combined.
Literal packaging and source-derived count stage
1,445 controls compare 469 tiny collections with every binary output word through length nine. Exact overlap DP matches their unrestricted optimum within this finite alphabet/domain. Paper forced counts do not exceed actual occurrences in every common word in that domain. Artifact round trips, original IDs, empty/duplicate/zero/UTF8 byte values, invalid models and budgets are checked.
python3 research/tools/superstring_packaging_reference.py research/prototypes/literal-packing-comparison-2026-10-09-v1/synthetic_greedy_gap-input.json
python3 research/tools/superstring_counts_reference.py research/prototypes/literal-packing-comparison-2026-10-09-v1/synthetic_periodic_windows-input.json
Seven saved comparisons and 31 emitted .ssp files include stable IDs/header/offset-length data. The public 48-title sample shrinks raw SSP1 by 28 bytes, leaves gzip unchanged at 1,685 and increases XZ from 1,616 to 1,648. A synthetic periodic sample shrinks SSP1 from 308 to 228 and gzip from 163 to 138, already with ordinary greedy. The generated greedy-gap control takes 10 raw bytes versus exact 9. Complete artifact objectives and API constraints matter; no customer, linker or access-performance result is inferred.
The count-stage reference implements source section02 only. For a^16 it recovers W=16 and m(a^k)=17-k. Source W validity in general depends on its unverified proof; finite comparisons do not certify that proof. The full forced-count/periodic-layer/request/link/cycle/Euler factor-two construction is unimplemented. The package model uses immutable length-aware bytes, not drop-in C strings or preserved pointer identity. Revised dossier.
Exact polynomial and numerical-range evidence
857 controls include 625 independent complex Hermitian 2-by-2 PSD oracle cases, zero pivots, exact rounding, eight generated matrices/polynomials, tensor order, finite Rayleigh probes, thresholds and budgets. An exact complete Gram PSD check supplies finite norm-threshold evidence independently of Crouzeix. Formal Python correctness was not proved.
python3 research/tools/numerical_range_polynomial_reference.py research/prototypes/numerical-range-comparison-2026-10-09-v1/nilpotent_sharp-input.json --grid 32
The separate twenty-direction PSD support enclosure and covered-cell Lipschitz/Frobenius calculation gives an entire-domain upper. The source constant-two result is conditional; the prior published constant is rounded upward to 483/200. On the sharp nilpotent example, the source-conditional upper about 2.189 fits tolerance 2.2 while the prior upper about 2.644 does not. A direct exact norm-two check already suffices with 231 charges, versus 84,404 for the full reference. No standalone library advantage is established. Eight-case report.
195 SciPy diagnostic controls reuse the existing isolated SciPy 1.18.1/NumPy 2.5.3 environment without installation changes. SVD/eigvalsh agreement is finite and tolerance bounded, not a rigorous library rounding guarantee. Revised numerical dossier.
Facility-planning evidence
5,127 controls include 4,992 models over all 52 three-point integer metrics with distances 1-4, compared with a separate all-subsets oracle. Rational primal/dual checks, false witnesses, full-metric conditions, model rejection, overlap and unknown budget states are included. The exact reference is conventional and exponential, not the source approximation algorithm.
python3 research/tools/metric_facility_reference.py research/prototypes/metric-facility-comparison-2026-10-09-v1/dual_tight_line-input.json --dual research/prototypes/metric-facility-comparison-2026-10-09-v1/dual_tight_line-dual.json --subset-budget 0
34 existing-solver controls compare eight generated cases with the preserved OR-Tools 9.15.6755 runtime. Seven OPTIMAL results match independent tiny optima. The 24-location/12-site case returns FEASIBLE with cost 12; a separate exact dual also gives 12 and proves optimality without enumeration. Four added controls, inputs and comparison.
Caps are 48 locations,128-bit canonical rational inputs, 100,000 enumerated subsets and five million enumeration distance lookups. Parsing/metric checks, primal/dual checks, rational bit arithmetic, allocations and time are outside those counters. Budget exhaustion preserves only completed evidence. Weights, capacities, directed distances and other real constraints need separate models. First product experiment.
Current documentary evidence
The twenty-one-contract registry now has 76 exact source-file pointers. The new metric-facility contract distinguishes fixed accuracy, strict rational metrics, unweighted demand, candidate budget, full source implementation, recovery promises and instance-specific bounds. Seven documentary controls passed. The new joined report has 16 gaps in a self-authored claim, 2,933 source hashes and six archived exact inputs; documentation completeness accepts no evidence contents.
Historical bundles and the frozen ten-case expert packet remain preserved. Source proof acceptance, actual reviewer comparison and paid customer value remain open. Feasibility, pilot.
Edit-distance source activation and existing exact tools
The selected source's large branch requires ell >= 2^1000. Its first necessary guard activates only at symbolic length 2^(2^(2^999))-1; full physical eligibility has additional accuracy and bit-width requirements. At N=10^12, ell=8. This makes the literal backend unsuitable for a practical first product. The source estimate is probabilistic on both sides and supplies no edit script.
A fresh isolated RapidFuzz 3.14.6 comparison adds 2,023 finite/library/guard controls. Exact routines already supply distances, cutoff predicates and alignments. On a generated 65,536-symbol shifted pair, unrestricted distance took about 101 ms median; a cutoff query took about 0.13 ms. These different output requirements and synthetic timings supply no source-backend speedup or customer benefit. A cutoff sentinel and normalization changes must be represented correctly. Detailed report, revised dossier.
Keep direct edit-distance commercialization deferred and conventional integration confidence low. The facility-planning review workflow remains the first product experiment; its buyer demand is still unvalidated.
Ordered-algebra formula testing
Family 116 adds an explicit characteristic-zero matrix construction and a bounded exact visible-tree implementation. A nonzero matrix entry disproves a free identity without the universal hitting theorem; a zero prefix is unknown and a full zero remains source-conditional. The source's black-box novelty and our visible-tree API are distinct. Positive-characteristic and inverse-formula generators have much larger dense/list costs and remain deferred.
726 controls pass across exact word expansion, matrix entries, existing SymPy, deliberately missed probes, target-algebra differences and resource limits. A source-derived short witness avoids full expansion on a generated sixteen-factor product, but conventional exact two-by-two probes find the same witness faster. Full source-dimension attempts hit rational-bit caps. No novel speed, accepted universal proof or paid demand is inferred. Dossier, comparison and schedules.
Retain a specialist SDK/CI integration experiment at low commercial confidence. The first product recommendation remains facility-planning assurance. An illustrative AUD 6,000/20-hour integration leaves AUD 2,400 at an assumed AUD 180/hour before other costs; 35 hours exceed the fee in labor alone. These are unvalidated assumptions.
Game algorithms and strategy-region evidence
Family 104 now has a four-variant matrix, bounded deterministic source reference and separate finite threshold-region checker. 1,097 controls pass across 368 generated games and boundary cases; every completed label agrees with the tiny exhaustive oracle. Some two-vertex recursions require hundreds of thousands of calls or exhaust two million charges. Mature production tools have not been benchmarked; no competitive advantage is shown.
The randomized huge-J wrapper is deferred when reached, but its Basic branch can bypass it. The stochastic binary discount denominator is not a required loop count; its backend remains unmeasured. Expected/almost-sure payoff, parity memory, value-optimality and initial credit require distinct contracts. Dossier, measurements.
Keep a specialist integration hypothesis at low confidence. An illustrative AUD 10,000/40-hour job leaves AUD 2,800 at AUD 180/hour before other costs; sixty hours exceeds the fee. These assumptions and demand remain unvalidated. Facility-planning assurance remains the first build.
First local facility-review release
The workbench now imports stable-ID CSVs, preserves exact decimals and raw bytes, rejects unsupported constraints, checks submitted and solver-proposed plans separately, verifies named-ID duals and emits fresh HTML/JSON reports. 104 new controls and eight generated examples pass. Cost 24 versus optimum 17 is a synthetic illustration; solver OPTIMAL can coexist with independent unknown, and a separate cost-12 dual example proves optimum without enumeration.
The mathematical source backend stays deferred. The tool accepts only the supported unweighted, uncapacitated strict metric, caps 48 locations and records resource limits. The site hosts generated reports; there is no customer upload endpoint or hosted optimizer. Browser intake, actual studies, measurable review benefit and paid demand remain open. Keep the existing price experiment explicitly unvalidated.
Evidence-package review and first-product fit
The family-266 review separates native exactly-three MUB evidence, an exact Fourier certificate and formal at-most-five/cancellation statements. Eleven named native package paths are absent from the pinned paper directory. Available exact Gram data pass 35 source-data/arithmetic controls, including exact factorization of six matrices and 98 positive pivots per order. The full moment expansion, native pipeline and Lean kernel were not run. Missing local artifacts are a reproduction gap, not a mathematical refutation. Existing execution/package tools constrain commercial differentiation.
The public planning workflow review reveals weighted, rectangular, required-site and capacity-assignment needs outside the initial facility checker. Prioritize a separate ordinary-plan audit before browser polish. Keep source metric eligibility distinct. The public baseline is synthetic and documented, not a customer dataset or a runtime benchmark. Demand and price hypotheses remain unvalidated.