Skip to content
Intermediate

A/B Test: Binary Design, Data Checks and Decision

Plan an independent binary-outcome experiment, check assignment counts and analyze the actual 2×2 table with effect intervals and explicit decision criteria.

product analystsdata scientists

Workflow

  1. Prespecify the decision and independent unit

    Define the user or other independent randomization unit, eligibility, binary outcome, observation window, allocation and practical improvement worth detecting. Set primary and guardrail metrics, alpha, the stopping rule and any multiplicity or sequential method before seeing outcomes. Repeated events from one assigned user are not automatically independent observations.

  2. Plan sample counts for the chosen design

    For a fixed-design equal-allocation binary test, use the baseline rate and an absolute percentage-point alternative in the sample planner. Check the resulting plan in the independent-binary-rate power mode using both sample counts and assumed rates. Plan exposure duration from eligible traffic and outcome maturity; continuous or dependent outcomes need a different design.

  3. Audit assignment and outcome completeness

    Compare actual assigned counts with the intended allocation using goodness of fit when its assumptions apply. Investigate sample-ratio mismatch, duplicate units, missing outcomes and logging changes before interpreting treatment effects. Build outcome denominators from the prespecified observation window and record every exclusion and unresolved data-quality issue.

  4. Analyze the independent binary outcome table

    Enter successes and failures for groups A and B in the separate 2×2 mode. Use the prespecified Pearson test without continuity correction when its approximation is suitable, or the specified Fisher conditional test for an appropriate sparse-table analysis. Report both rates, B-minus-A percentage-point difference and the explicitly labeled Newcombe interval.

  5. Interpret uncertainty with the experiment context

    Compare the test p-value with prespecified alpha while retaining its actual precision. Evaluate effect size, interval, guardrails, implementation cost and data limitations together. The displayed Newcombe interval is not the inversion of the Fisher test, and threshold crossing alone is not evidence that an experiment should ship.

  6. Record the decision and follow-up evidence

    Choose release, further study or no change using the original decision criteria and observed evidence. Save allocation counts, outcome table, test/interval methods, exclusions, stopping rule and any deviations. If the study used repeated peeking, clustering or multiple tests, use the corresponding analysis rather than relabeling this fixed-design calculator result.

Tools Used

Checklist

0 / 6 completed

Loading your checklist…

Prespecify the decision and independent unit

Plan sample counts for the chosen design

Audit assignment and outcome completeness

Analyze the independent binary outcome table

Interpret uncertainty with the experiment context

Record the decision and follow-up evidence

Reference Materials

Categorical goodness of fitStandard

SciPy documents goodness of fit for category counts and the need for compatible totals and justified degrees of freedom. An allocation check is a different question from the treatment-outcome table.

Fisher’s conditional testStandard

SciPy documents the two-sided conditional table-probability definition used by the Fisher mode. The difference interval is calculated separately.

P-values in contextStandard

The ASA statement emphasizes transparent reporting and distinguishes statistical thresholds from effect size and the importance of a result.

Experiment decision recordTable

Keep the following evidence with the actual version used for this task.

RecordIncludeCheck
DesignUnit, outcome, alternative, allocation and stop rulePrespecified before outcomes
EvidenceAssignment counts, 2×2 outcomes and exclusionsInvestigate data-quality problems
DecisionRates, difference, interval, test and guardrailsKeep practical costs and limitations visible
  • Keep assignment and outcome tables separate

    A sample-ratio check does not test whether treatment changed conversion.

  • Report an inconclusive result honestly

    A nonsignificant test is not proof that effects are absent or variants are equivalent.