# First product: Optimization Assurance Workbench

Decision dated 9 October 2026, edition v3. The best current product hypothesis is a focused facility-planning review tool for logistics and operations consultancies. This is the first experiment within the shared evidence platform, with no validated demand, paid customer or proprietary algorithm advantage. [Source and technical dossier](../opportunities/011-facility-placement-planner/2026-10-09-v4.md).

## Buyer, job and deliverable

A consultancy preparing a depot or service-site study imports its data and proposed facilities. The workbench checks the declared model, reproduces the cost, compares alternatives through an established solver, and produces a report suitable for the consultancy's own review. The proposed benefit is fewer unsupported conclusions and less repetitive checking across study revisions. No interviews have established how often this problem occurs or whether existing tools already solve it.

The first paid deliverable would be an assisted review of one clearly scoped study. A subscription becomes a separate hypothesis after repeated use, measurable time savings and manageable support effort exist. The mathematical collection supplies research and assumption-review material; the immediate checks use conventional exact arithmetic and optimization.

## First software release

| Step | User-facing behavior | Acceptance condition |
| --- | --- | --- |
| Import | Upload or locally select stable site/client IDs, a distance matrix, selected sites and units | Reject duplicate IDs and missing cells; preserve the original bytes and mapping |
| Declare model | Record demand weights, facility budget, capacities, mandatory sites and distance provenance | Every supplied business constraint is supported, explicitly unresolved or rejected; never silently drop it |
| Check scope | Explain which mathematical assumptions hold on the frozen data | A failed strict-metric test does not declare an ordinary directed planning problem invalid |
| Check plan | Recompute assignments, feasibility and cost independently | Use original data and exact supported arithmetic; separate measurement uncertainty |
| Compare | Run a recorded conventional solver and compare proposed plans | Preserve time limit, status, model version, candidate IDs and a separately checked witness |
| Establish quality | Check an optional rational lower-bound certificate or complete a bounded reference | Show a certified gap only when both bounds are justified; show unknown when they are not |
| Export | Produce a readable report and machine-readable evidence bundle | Include original input hashes, versions, assumptions, completed checks, unresolved items and review sign-off |

Current implementation: a local CSV-to-report workflow with stable IDs, exact decimal/rational import, explicit model checks, separately checked submitted/solver plans, optional named-ID dual certificates and fresh HTML/JSON evidence exports. Eight generated reviews and 104 new controls demonstrate this release. [Open examples and run instructions](../prototypes/facility-review-workbench-2026-10-09-v2/report.md). The website hosts those synthetic reports as static files. It has no customer upload endpoint, accounts or hosted solver. Browser intake, further model variants, automated dual generation and production deployment remain future work.

## The concrete demonstration

On a generated line with 24 locations and a budget of 12 sites, the two-second existing-solver run returned FEASIBLE with a checked cost of 12. The independent subset reference deliberately stopped after ten of 2,704,156 possible k-subsets. A separate hand-constructed dual with alpha=1 at every client, beta=1 on coincident client/site pairs and lambda=1 gives lower bound 24-12=12. Its inequalities are checked exactly, so that candidate is independently optimal for the declared model.

This illustrates useful evidence that can survive a solver time limit. It is one easy synthetic instance, with a specially constructed certificate, and does not show superiority over a longer solver run or existing expert practice. [Inputs, statuses and results](../prototypes/metric-facility-comparison-2026-10-09-v1/report.md).

## Implementation order

1. The local review/export workflow now runs. First add a separately scoped weighted, rectangular cost/plan audit, with explicit required-site and assignment semantics. The public PySAL workflow exposes these gaps; browser polish alone does not address them.
2. The optional OR-Tools adapter now preserves exact scaling, records limits/status and independently checks returned plans. Benchmark relevant permissioned studies and retain solver best bounds as reported until separately certified.
3. Add weighted demand and capacity checking only as distinct, tested model variants. Source-theorem eligibility stays separate. Use permissioned study examples to prioritize which variant matters first.
4. Add database, accounts and hosted uploads after a real recurring workflow and data-handling requirements are established. Preserve the static research catalog alongside the product.
5. Revisit the repository's new approximation engine after source proof acceptance, numerical parameter selection and representative performance measurements. Its asymptotic guarantee alone supplies no production advantage.

The reference currently caps 48 locations and exact 128-bit rational inputs. Full triangle checking is cubic; exhaustive subset search is capped and exponential. Those are a starting evidence module, not production-scale capacity promises. General LP-dual extraction, rational repair of floating solutions, solver-proof authentication and physical-data uncertainty remain separate engineering work.

## Price and delivery experiment

Hypothesized assisted-review price: AUD 5,000-10,000. At AUD 7,500 and an assumed 25 hours at AUD 180/hour, labor costs AUD 4,500, leaving AUD 3,000 before overhead and other costs. At 45 hours, labor alone costs AUD 8,100. These assumptions must be replaced with actual delivery measurements and purchase evidence.

The initial evaluation should measure whether a consultant has repeated studies, material review gaps and budget for independent checks. Prepare interview and pilot materials, but contact nobody without explicit outreach authorization. A useful trial should record the original review process, independently adjudicated errors, false alarms, preparation/review time, decisions changed, support time and willingness to pay. Synthetic scope-error fixtures cannot measure these outcomes.

## Stop or change direction

Change the plan if existing solver reports and in-house review already meet the need, customers require unsupported constraints before any value is delivered, the checks find no actionable issues, or achievable fees cannot cover delivery. Do not claim revenue, accuracy savings or novelty based on the number of research papers or finite tests.

The first release should be assessed against established alternatives, including [PySAL spopt's p-median workflow](https://pysal.org/spopt/notebooks/p-median.html) and [OR-Tools' documented statuses and integer model](https://developers.google.com/optimization/cp/cp_solver), reviewed 9 October 2026. Building and reporting on an optimization model is already possible; the proposed paid value is a demonstrably better review workflow.

## First local release evidence

The six-location example improves submitted cost 24 to exact optimum 17. A separate run reports solver OPTIMAL at 17 while independent optimality stays unknown when enumeration is disabled. The named-ID 24-location certificate proves cost 12 optimal without enumeration. All are generated controls, not customer savings. Unsupported capacities and directed distances receive explicit outcomes. Original byte archives and report histories are preserved; no source approximation engine or profitable demand is claimed.

## Public workflow fit review

The [documented planning comparison](facility-workflow-fit-2026-10-09-v1.md) changes the next implementation priority: ordinary weighted/rectangular plan checks come before browser polish. Capacities require explicit assignments and load checks; the strict metric module must retain its original contract. This is a public synthetic baseline, not customer validation or a PySAL runtime comparison.
