# Fixed-margin data sandbox and sampling-law audit

Edition: 9 October 2026 Australia/Brisbane. Decision: **Prototype the law/reference workflow; defer the literal research backend**. Demand and profitability remain unvalidated.

## Research finding

Family 115 claims exact uniform sampling of unbounded nonnegative integer tables with arbitrary equal-total margins in expected polynomial bit complexity, and bounded-time approximate sampling. A separate paper claims an FPRAS for counting cell-bounded tables, including structural zeros. The selected formal scope matches these distinctions; it has not been independently checked here. Jointly variable dimensions and binary margin sizes are material theoretical improvements over restricted earlier regimes.

## Problem and buyer

A data-testing or statistical-model team needs aggregate scenarios that preserve exact margins and a clear statement of the probability law. Merely preserving totals does not make a sampler uniform over integer tables. Conversely, uniform tables may be the wrong law for an independence test. A buyer could value constraint fidelity and an auditable choice of law inside a repeated test or model-review workflow.

## What the finding could enable

If practical constants are improved, the source construction could extend uniform-table generation to unrestricted margins and dimensions with binary-size complexity. That backend is not implemented. The current deliverable is a bounded exact reference and sampling-law report: count feasible tables, replay a table by rank, respect cell bounds, expose infeasibility or unknown budget status, and compare intended laws. It could introduce an aggregate-scenario testing plugin or a model-assurance integration. Its conventional finite dynamic program is a benchmark component rather than the new unrestricted algorithm.

## Technical and commercial limits

The declared general outer sampler runs `d^200(k+b)^2` transitions per trial, with `d=10+(m+1)(n+1)`; the two-by-two example already has a 260-digit transition count. Exact correction uses a rare exhaustive branch, so expected polynomial time is distinct from a deadline guarantee on every run. The cell-bounded counting paper explicitly prioritizes complexity over practical runtime; its headline does not itself supply the unbounded paper's exact sampler with structural zeros.

Uniform tables also differ from the conditional-independence law. For margins `[2,2]` in both directions, uniform probabilities are `1/3,1/3,1/3`, while conditional independence gives `1/6,2/3,1/6`. The reference is capped at dimension six, total 128 and work/state budgets. Exact margins and either law provide no privacy, realism or relational-data guarantee. Do not use a uniform-table law as an automatic substitute in existing independence tests.

## Minimal architecture

Margin/bound validation -> explicit probability-law contract -> finite exact reference or existing statistical backend -> recorded rank/seed and arithmetic model -> constraint checks -> law comparison -> export. Keep a future research sampler behind a separate capability flag with its own implementation, runtime and distribution evidence. Budget exhaustion remains unknown rather than being reported as zero feasible tables.

## Existing alternatives and differentiation

SciPy `random_table` supplies fixed-margin conditional-independence sampling using Boyett/Patefield methods; R `r2dtable` also uses Patefield. Tonic addresses broader data generation/deidentification. The proposed differentiation is law and constraint evidence within one aggregate testing workflow, or a demonstrated new backend crossover if achieved. It is not a claim that those tools implement the wrong distribution or lack all needed features. [SciPy API](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.random_table.html), [R API](https://www.stat.math.ethz.ch/R-manual/R-patched/RHOME/library/stats/html/r2dtable.html), [Tonic products](https://www.tonic.ai/pricing).

## Monetization hypothesis

Hypothesis: AUD 3,000–6,000 for a bounded law/constraint integration assessment. At an assumed AUD 5,000 fee and 20 specialist hours at AUD 180/hour, AUD 1,400 remains before sales, support and overhead. A subscription is credible only for a repeated workflow with measured value and low support effort. The small reference script can be open; paid value would need to come from integration, evidence and support. These are experimental prices, not quotes or revenue.

## Validation experiment

The current reference passed 325 finite checks against two-by-two closed forms, independent Cartesian enumeration, permutation counts, structural zeros and unknown-budget behavior. Exact rank bijections verify small state-space coverage; they do not validate the source's general sampler. The saved law contrast has total-variation distance `1/3` for the two-by-two example. A new exact event-law audit now uses public R/SciPy documentation examples. It holds the declared event fixed across the intended and candidate laws, and explicitly conditions independence when bounds are supplied. In the public [[8,2],[1,5]] example the probabilities 5/143 versus 3/7 cross the declared 1/20 threshold. This is a model-substitution consequence, not an existing-tool error. 617 controls cover closed forms, labelled assignments, bounds and unknown behavior. A fresh isolated SciPy 1.18.1/NumPy 2.5.3 comparison now passes 586 controls on 284 margin pairs/494 observed tables and five public/generated fixtures. Earlier failed 1.15.3 environments remain preserved; R was not run. [Concrete report](../../prototypes/public-contingency-law-review-2026-10-09-v2.md). Next compare an actual aggregate-test workflow, holding intended law, event and constraints fixed. Measure preparation effort, errors found and delivery cost separately from theoretical backend runtime.

## Conditions to reject or defer

Defer the literal unrestricted research backend until substantially better practical parameters or a measured restricted regime exist. Reject the workflow business if teams already choose/document laws correctly, do not need exact aggregate constraints, or integration effort consumes the fee. Defer privacy claims without an independent privacy mechanism. A wrong-law comparison is not evidence of a runtime speedup.

## Next concrete action

Run [the public law/event report](../../prototypes/public-contingency-law-review-2026-10-09-v2.md) and [audit_contingency_law.py](../../tools/audit_contingency_law.py). Choose a realistic recurring aggregate workflow and measure errors found, preparation/support cost and whether explicit constraints actually model the task. The finite live-library comparison is complete; large-support scaling, actual experiment design and recurring buyer value remain unmeasured. Continue searching for practical restricted sampler parameters without substituting the reference for the source theorem.

## Pinned research sources

- [Uniform sampler manuscript](https://github.com/openai/math/blob/fd4aeeb2ee4fc729c18d98444fed42fd0529eeeb/preprints/Exact-Uniform-Sampling-of-Contingency-Tables-with-Arbitrary-Margins-September-24-2026/main.pdf): introduction, parameters, outer schedule, law tabulation and exact correction read.
- [Cell-bounded counter manuscript](https://github.com/openai/math/blob/fd4aeeb2ee4fc729c18d98444fed42fd0529eeeb/preprints/An-FPRAS-for-Cell-Bounded-Contingency-Tables-September-24-2026/main.pdf): introduction, preprocessing/scales and selected algorithm schedule read; full proof review remains pending.
- [Selected formal scope](https://github.com/openai/math/blob/fd4aeeb2ee4fc729c18d98444fed42fd0529eeeb/lean/docs/115.md).
- [Explicit feasibility analysis](../../FEASIBILITY-2026-10-09-v2.md) and [saved schedule calculations](../../snapshots/2026-10-08-baseline/direct-candidate-parameters-2026-10-09-v1.json).

## New flow consequence

The cell-bounded paper also counts and approximately samples bounded integer arc-value flows through a private-vertex table bijection. This is a separate network-scenario hypothesis, with different sampling accuracy/runtime and selected formal scope. [Dossier 041](../041-bounded-flow-scenario-analysis/2026-10-09-v1.md). Neither the exact uncapped table sampler nor a conditional-independence law transfers silently to bounded uniform flows.
