A/B Test Design Workbench

Plan feature-launch experiments across metric types, power, duration, multiple testing risk, A/A checks, and CUPED variance reduction.

Experimentation
Intermediate
Experiment Scenario

The test math changes depending on whether the metric is a proportion or an average.

10%

Current conversion rate before launching the new feature.

5%

Minimum meaningful change worth detecting.

5%

False-positive rate for one comparison.

80%

Beta is 20% at this power target.

Traffic & Multiple Testing
10,000
2
100%
3
4
0.40

Higher correlation means pre-experiment data can remove more noise.

Decision Summary
What this launch test needs before you can trust the read.
Expected Effect
0.5 pts / 5.0%
absolute and relative lift
Base Sample Size
57,645
per variant before multiplicity correction
Adjusted Alpha
0.42%
12 planned comparisons
Adjusted Duration
21 days
5,000 users / variant / day
False Positive Risk
46.0%
chance of at least one false positive
With CUPED
17 days
84,546 users / variant
How to Read This
The calculator separates a single clean test from the messier reality of many metrics and segments.

For a 0.5 pts / 5.0% effect at 5% alpha and 80% power, the clean two-variant test needs 57,645 users per variant.

Because you are checking 12 comparisons, the chance of at least one false positive rises to 46.0%. Bonferroni keeps the family-wise error rate controlled by using 0.42% alpha per comparison.

CUPED reduces noise when pre-experiment behavior predicts the metric. With correlation 0.40, it saves about 4 days after the multiple-testing adjustment.

7, 14, 28, or Longer?
Fixed run lengths trade speed against power, seasonality, and smaller detectable effects.

7 days

59% power
Sample / variant
35,000
Detectable effect
0.6 pts
Relative MDE
6.3%

Fast, but usually underpowered and noisy.

14 days

87% power
Sample / variant
70,000
Detectable effect
0.4 pts
Relative MDE
4.5%

Often a practical minimum with two weekly cycles.

28 days

99% power
Sample / variant
140,000
Detectable effect
0.3 pts
Relative MDE
3.2%

More stable read with seasonality and novelty risk reduced.

56 days

100% power
Sample / variant
280,000
Detectable effect
0.2 pts
Relative MDE
2.2%

Best for small effects or low-traffic segments, but slower learning.

Power Curve
How sample size changes your chance of detecting the expected effect.
Multiple Testing Risk
More metrics and segments increase the chance of at least one false positive.
CUPED Savings
Pre-experiment data can reduce variance when it predicts the experiment metric.
After multiple-testing adjustment100,649 / variant
With CUPED84,546 / variant

Correlation of 0.40 reduces variance by 16%, saving about 4 days at current traffic.

A/A Testing Readiness
A/A tests should mostly find nothing. Their job is to catch instrumentation and assignment problems before a real launch.
10

Even with no real effect, repeated checks can produce false positives.

Expected false positives

0.50

At least one false positive

40.1%

  • Instrumentation is stable and logging every assignment.
  • Randomization happens before users can self-select into variants.
  • Metric definitions are frozen before the test starts.
  • Eligibility, exposure, and analysis populations are clearly defined.
  • The team agrees not to keep slicing until something looks significant.