On this page
Fixed-margin data sandbox and sampling-law audit
Edition: 9 October 2026 Australia/Brisbane. Decision: Prototype the law/reference workflow; defer the literal research backend. Demand and profitability remain unvalidated.
Research finding
Family 115 claims exact uniform sampling of unbounded nonnegative integer tables with arbitrary equal-total margins in expected polynomial bit complexity, and bounded-time approximate sampling. A separate paper claims an FPRAS for counting cell-bounded tables, including structural zeros. The selected formal scope matches these distinctions; it has not been independently checked here. Jointly variable dimensions and binary margin sizes are material theoretical improvements over restricted earlier regimes.
Problem and buyer
A data-testing or statistical-model team needs aggregate scenarios that preserve exact margins and a clear statement of the probability law. Merely preserving totals does not make a sampler uniform over integer tables. Conversely, uniform tables may be the wrong law for an independence test. A buyer could value constraint fidelity and an auditable choice of law inside a repeated test or model-review workflow.
What the finding could enable
If practical constants are improved, the source construction could extend uniform-table generation to unrestricted margins and dimensions with binary-size complexity. That backend is not implemented. The current deliverable is a bounded exact reference and sampling-law report: count feasible tables, replay a table by rank, respect cell bounds, expose infeasibility or unknown budget status, and compare intended laws. It could introduce an aggregate-scenario testing plugin or a model-assurance integration. Its conventional finite dynamic program is a benchmark component rather than the new unrestricted algorithm.
Technical and commercial limits
The declared general outer sampler runs d^200(k+b)^2 transitions per trial, with d=10+(m+1)(n+1); the two-by-two example already has a 260-digit transition count. Exact correction uses a rare exhaustive branch, so expected polynomial time is distinct from a deadline guarantee on every run. The cell-bounded counting paper explicitly prioritizes complexity over practical runtime; its headline does not itself supply the unbounded paper's exact sampler with structural zeros.
Uniform tables also differ from the conditional-independence law. For margins [2,2] in both directions, uniform probabilities are 1/3,1/3,1/3, while conditional independence gives 1/6,2/3,1/6. The reference is capped at dimension six, total 128 and work/state budgets. Exact margins and either law provide no privacy, realism or relational-data guarantee. Do not use a uniform-table law as an automatic substitute in existing independence tests.
Minimal architecture
Margin/bound validation -> explicit probability-law contract -> finite exact reference or existing statistical backend -> recorded rank/seed and arithmetic model -> constraint checks -> law comparison -> export. Keep a future research sampler behind a separate capability flag with its own implementation, runtime and distribution evidence. Budget exhaustion remains unknown rather than being reported as zero feasible tables.
Existing alternatives and differentiation
SciPy random_table supplies fixed-margin conditional-independence sampling using Boyett/Patefield methods; R r2dtable also uses Patefield. Tonic addresses broader data generation/deidentification. The proposed differentiation is law and constraint evidence within one aggregate testing workflow, or a demonstrated new backend crossover if achieved. It is not a claim that those tools implement the wrong distribution or lack all needed features. SciPy API, R API, Tonic products.
Monetization hypothesis
Hypothesis: AUD 3,000–6,000 for a bounded law/constraint integration assessment. At an assumed AUD 5,000 fee and 20 specialist hours at AUD 180/hour, AUD 1,400 remains before sales, support and overhead. A subscription is credible only for a repeated workflow with measured value and low support effort. The small reference script can be open; paid value would need to come from integration, evidence and support. These are experimental prices, not quotes or revenue.
Validation experiment
The current reference passed 325 finite checks against two-by-two closed forms, independent Cartesian enumeration, permutation counts, structural zeros and unknown-budget behavior. Exact rank bijections verify small state-space coverage; they do not validate the source's general sampler. The saved law contrast has total-variation distance 1/3 for the two-by-two example. A new exact event-law audit now uses public R/SciPy documentation examples. It holds the declared event fixed across the intended and candidate laws, and explicitly conditions independence when bounds are supplied. In the public [[8,2],[1,5]] example the probabilities 5/143 versus 3/7 cross the declared 1/20 threshold. This is a model-substitution consequence, not an existing-tool error. 617 controls cover closed forms, labelled assignments, bounds and unknown behavior. A fresh isolated SciPy 1.18.1/NumPy 2.5.3 comparison now passes 586 controls on 284 margin pairs/494 observed tables and five public/generated fixtures. Earlier failed 1.15.3 environments remain preserved; R was not run. Concrete report. Next compare an actual aggregate-test workflow, holding intended law, event and constraints fixed. Measure preparation effort, errors found and delivery cost separately from theoretical backend runtime.
Conditions to reject or defer
Defer the literal unrestricted research backend until substantially better practical parameters or a measured restricted regime exist. Reject the workflow business if teams already choose/document laws correctly, do not need exact aggregate constraints, or integration effort consumes the fee. Defer privacy claims without an independent privacy mechanism. A wrong-law comparison is not evidence of a runtime speedup.
Next concrete action
Run the public law/event report and audit_contingency_law.py. Choose a realistic recurring aggregate workflow and measure errors found, preparation/support cost and whether explicit constraints actually model the task. The finite live-library comparison is complete; large-support scaling, actual experiment design and recurring buyer value remain unmeasured. Continue searching for practical restricted sampler parameters without substituting the reference for the source theorem.
Pinned research sources
- Uniform sampler manuscript: introduction, parameters, outer schedule, law tabulation and exact correction read.
- Cell-bounded counter manuscript: introduction, preprocessing/scales and selected algorithm schedule read; full proof review remains pending.
- Selected formal scope.
- Explicit feasibility analysis and saved schedule calculations.
New flow consequence
The cell-bounded paper also counts and approximately samples bounded integer arc-value flows through a private-vertex table bijection. This is a separate network-scenario hypothesis, with different sampling accuracy/runtime and selected formal scope. Dossier 041. Neither the exact uncapped table sampler nor a conditional-independence law transfers silently to bounded uniform flows.