# First product: Optimization Assurance Workbench

Decision dated 9 October 2026. The best current product hypothesis is a focused facility-planning review tool for logistics and operations consultancies. This is the first experiment within the shared evidence platform, with no validated demand, paid customer or proprietary algorithm advantage. [Source and technical dossier](../opportunities/011-facility-placement-planner/2026-10-09-v2.md).

## Buyer, job and deliverable

A consultancy preparing a depot or service-site study imports its data and proposed facilities. The workbench checks the declared model, reproduces the cost, compares alternatives through an established solver, and produces a report suitable for the consultancy's own review. The proposed benefit is fewer unsupported conclusions and less repetitive checking across study revisions. No interviews have established how often this problem occurs or whether existing tools already solve it.

The first paid deliverable would be an assisted review of one clearly scoped study. A subscription becomes a separate hypothesis after repeated use, measurable time savings and manageable support effort exist. The mathematical collection supplies research and assumption-review material; the immediate checks use conventional exact arithmetic and optimization.

## First software release

| Step | User-facing behavior | Acceptance condition |
| --- | --- | --- |
| Import | Upload or locally select stable site/client IDs, a distance matrix, selected sites and units | Reject duplicate IDs and missing cells; preserve the original bytes and mapping |
| Declare model | Record demand weights, facility budget, capacities, mandatory sites and distance provenance | Every supplied business constraint is supported, explicitly unresolved or rejected; never silently drop it |
| Check scope | Explain which mathematical assumptions hold on the frozen data | A failed strict-metric test does not declare an ordinary directed planning problem invalid |
| Check plan | Recompute assignments, feasibility and cost independently | Use original data and exact supported arithmetic; separate measurement uncertainty |
| Compare | Run a recorded conventional solver and compare proposed plans | Preserve time limit, status, model version, candidate IDs and a separately checked witness |
| Establish quality | Check an optional rational lower-bound certificate or complete a bounded reference | Show a certified gap only when both bounds are justified; show unknown when they are not |
| Export | Produce a readable report and machine-readable evidence bundle | Include original input hashes, versions, assumptions, completed checks, unresolved items and review sign-off |

Current implementation: a local command-line reference for strict, unweighted, uncapacitated rational metrics; generated existing-solver comparisons; exact primal/dual checking; versioned research artifacts. The live website is a static catalog. It does not accept customer data or execute uploaded files. CSV intake, a review interface, production solver integration and automated general dual generation are planned work.

## The concrete demonstration

On a generated line with 24 locations and a budget of 12 sites, the two-second existing-solver run returned FEASIBLE with a checked cost of 12. The independent subset reference deliberately stopped after ten of 2,704,156 possible k-subsets. A separate hand-constructed dual with alpha=1 at every client, beta=1 on coincident client/site pairs and lambda=1 gives lower bound 24-12=12. Its inequalities are checked exactly, so that candidate is independently optimal for the declared model.

This illustrates useful evidence that can survive a solver time limit. It is one easy synthetic instance, with a specially constructed certificate, and does not show superiority over a longer solver run or existing expert practice. [Inputs, statuses and results](../prototypes/metric-facility-comparison-2026-10-09-v1/report.md).

## Implementation order

1. Wrap the current reference in a local review workflow with stable IDs, explicit assumptions and an exportable report. Keep parsing/model adapters separate from the exact checker.
2. Implement a reusable conventional solver adapter, preserving exact rational scaling where supported. Independently check returned plans. Treat solver best bounds as reported until separately certified.
3. Add weighted demand and capacity checking only as distinct, tested model variants. Source-theorem eligibility stays separate. Use permissioned study examples to prioritize which variant matters first.
4. Add database, accounts and hosted uploads after a real recurring workflow and data-handling requirements are established. Preserve the static research catalog alongside the product.
5. Revisit the repository's new approximation engine after source proof acceptance, numerical parameter selection and representative performance measurements. Its asymptotic guarantee alone supplies no production advantage.

The reference currently caps 48 locations and exact 128-bit rational inputs. Full triangle checking is cubic; exhaustive subset search is capped and exponential. Those are a starting evidence module, not production-scale capacity promises. General LP-dual extraction, rational repair of floating solutions, solver-proof authentication and physical-data uncertainty remain separate engineering work.

## Price and delivery experiment

Hypothesized assisted-review price: AUD 5,000-10,000. At AUD 7,500 and an assumed 25 hours at AUD 180/hour, labor costs AUD 4,500, leaving AUD 3,000 before overhead and other costs. At 45 hours, labor alone costs AUD 8,100. These assumptions must be replaced with actual delivery measurements and purchase evidence.

The initial evaluation should measure whether a consultant has repeated studies, material review gaps and budget for independent checks. Prepare interview and pilot materials, but contact nobody without explicit outreach authorization. A useful trial should record the original review process, independently adjudicated errors, false alarms, preparation/review time, decisions changed, support time and willingness to pay. Synthetic scope-error fixtures cannot measure these outcomes.

## Stop or change direction

Change the plan if existing solver reports and in-house review already meet the need, customers require unsupported constraints before any value is delivered, the checks find no actionable issues, or achievable fees cannot cover delivery. Do not claim revenue, accuracy savings or novelty based on the number of research papers or finite tests.

The first release should be assessed against established alternatives, including [PySAL spopt's p-median workflow](https://pysal.org/spopt/notebooks/p-median.html) and [OR-Tools' documented statuses and integer model](https://developers.google.com/optimization/cp/cp_solver), reviewed 9 October 2026. Building and reporting on an optimization model is already possible; the proposed paid value is a demonstrably better review workflow.
