Minimum Detectable Effect Calculator

Minimum Detectable Effect Calculator

Estimate the smallest effect your design can reliably detect. Switch between continuous means and binary proportions, then enter sample size, alpha, power, allocation, standard deviations, or baseline rate to see the MDE and formula steps.

🎯MDE Planning Presets

📝Design Inputs

Means use SD inputs. Proportions use baseline p.

Two-sided designs use z alpha/2.

Lower alpha makes the MDE larger.

Power = 1 - beta, the detection probability at the MDE.

Control, baseline, or reference group count.

Treatment, variant, or comparison group count.

Use the planning SD for the reference group.

Use the planning SD for the comparison group.

Minimum detectable effect 0 outcome units
Detectable endpoint 0 comparison value at MDE
Standard error 0 before z multiplier
Standardized effect 0 Cohen d or h

🔢Current Design Snapshot

1.960z alpha
0.842z beta
1.00n2 / n1
10000total n

⚙Formula Breakdown

Means MDEdelta = (zα + zβ) x sqrt(sigma1²/n1 + sigma2²/n2). This is the standard normal planning formula for two independent means.
Proportion MDEApproximate MDE = (zα + zβ) x sqrt(p(1 - p)(1/n1 + 1/n2)). The calculator reports it as percentage points from the baseline p.
Critical valueFor two-sided tests, zα = inverse normal(1 - alpha / 2). For one-sided tests, zα = inverse normal(1 - alpha).
Power valuezβ = inverse normal(power). Common values are 0.842 for 80% power, 1.282 for 90% power, and 1.645 for 95% power.
Allocation effectThe calculator uses the entered n1 and n2 directly. Unequal group sizes increase the standard error when the smaller group becomes thin.
Power checkFor proportions, the headline MDE uses the common baseline-variance approximation and then checks achieved power with the alternative endpoint variance.

📊Z Values and Planning Choices

SettingAlpha or BetaTwo-Sided zOne-Sided zPower zMDE Impact
90% confidencealpha 0.101.6451.282variesSmaller MDE
95% confidencealpha 0.051.9601.645variesStandard default
99% confidencealpha 0.012.5762.326variesLarger MDE
80% powerbeta 0.20same alphasame alpha0.842Common screen
90% powerbeta 0.10same alphasame alpha1.282Stricter plan
95% powerbeta 0.05same alphasame alpha1.645Largest MDE

đź§ŞBaseline Proportion Reference

Baseline pp(1 - p)Relative VarianceWhere SeenPlanning Note
1%0.0099LowRare eventsNormal check may be needed
5%0.0475ModerateDefects, clicksSmall pp lifts need large n
10%0.0900HigherConversion testsCommon product baseline
25%0.1875HighAdoption ratesMDE grows near the middle
50%0.2500HighestSurvey sharesMaximum binomial variance
75%0.1875HighRetention ratesSame variance as 25%
90%0.0900HigherPass ratesMirror of 10%

âš–Allocation Ratio Reference

n2 / n1Group 2 ShareSE PatternWhen It FitsCaution
0.2520%Much widerScarce variant exposureThin group drives MDE
0.5033%WiderLimited treatment capacityTotal n often rises
1.0050%Most efficientEqual cost observationsBest default
1.5060%Slightly widerMore variant trafficUsually acceptable
2.0067%WiderRollout preferenceControl gets thin
4.0080%Much widerRare control plansCheck total n carefully

🔍MDE Scenario Comparison Grid

ScenarioOutcomen1n2Input 1Input 2AlphaPowerTailRead
A/B conversionProportion5000500010% pbaseline0.0580%GreaterSmall lift test
Email click rateProportion12000120004% pbaseline0.0590%GreaterRare event lift
Defect reductionProportion9009008% pbaseline0.0590%LessQuality drop
Survey supportProportion70070050% pbaseline0.0580%TwoHighest variance
Blood pressureMean909012 SD12 SD0.0580%TwoClinical units
Website latencyMean26026090 SD85 SD0.0590%LessMilliseconds
Test scoresMean1601609 SD10 SD0.0580%TwoScale points
Battery lifeMean75752.5 SD2.8 SD0.0590%TwoHours
Uneven rolloutProportion4000800018% pbaseline0.0580%Greater2:1 split

đź’ˇMDE Planning Tips

Use the smallest useful effect: MDE is a planning threshold, not a promise. If your design can only detect a very large lift, a null result will not rule out smaller useful effects.
Keep units straight: For means, sigma and MDE are in the outcome's natural units. For proportions, the calculator reports absolute percentage points from the baseline rate.
Watch imbalance: A 50/50 split is usually most efficient. If one group is much smaller, the standard error grows even when total sample size looks large.
Check assumptions: Normal approximations work best with adequate sample size, independent observations, stable variance, and enough successes and failures for proportion outcomes.

So first: all experiments start with a question about capability rather then outcome. Before you run the test, you want to know what the design can realy see. Too faint? Then the noise swamps the signal, and you get a null result; meaning nothing. The minimum detectable effect answer a practical question. What’s the smallest difference your study could reliably find under these circumstances? In other words: The resolution of your lens. Maybe you have a blurry lens that sees general shape of a big tree, but doesn’t see individual flowers down by the bottom. But you get to choose what that means.

Once you plug in your variance and your sample size, the calculator do all the math for you. So you don’t waste time trying to convert and guess at coefficients. First, you enter in your baseline value. This matter more than you would think. Variance reaches its peak in the middle for proportions. In other words, if you’re measuring something with an even split between two adult-sized sofa (like, say, support for a given political candidate), then you need more respondents because your data’s inherently noisy. On the other hand, if you’re looking at something rarer, like defects, then you see the same effect size with fewer respondents. The page has a nice reference table laying this out. You can see how the math changes based off moving from one percent to, say, fifty.

What Is Minimum Detectable Effect?

There’s another tradeoff here as well: sensitivity vs. Certainty. Alpha means you don’t want false alarms, which is thinking you see something when it is just a ghost. Power means you don’t want to miss something real, like a true signal. People typically sets their power at 80% and their alpha at 5%. It is not a law but a convention. Bump up your power to 90%, and you’re insisting on more evidence. To cross the higher bar you set, you need a louder effect, or a bigger sample size. So the MDE get bigger.

“Everyone pulls the lever of sample size. The more people you have, the smaller the standard error. That makes the MDE tighter. But it’s not linear. To half the detectable effect, you need a sample that’s four times larger. And that’s a brutal bit of geometry, one reason so many tests flop. Without understanding this, teams assume tiny lifts and don’t realize they must gather an army-sized set of data. The tool makes this reality clear in seconds. It translates abstract goals into concrete numbers.”

Another silent power-killer is allocation. Wasting precision by splitting traffic ineffectively sucks. You can’t have half go to the variant and eight-tenths to the control; that’s only going to increase your standard error. That small group will be the bottleneck, it will determine how certain you are. The reason a fifty-fifty split is so efficient isn’t an accident. It perfectly balances each group’s spread. Straying from that mean you need a far bigger overall sample to make up for it.

What the calculator doesn’t do is give you context. The MDE is a stat limit, not a business limit. Does the effect you’re detecting matter? Are you lifting conversion by half a percent and you have a million users? That’s statistically significant, but it won’t pay for the feature. That’s a different question. The MDE answers “what can I find”. It does not answer “what should I care about.

Define the lowest-power effect you could imagine changing your mind about. Then work backwards. Can your current tools finds this? Examine your inputs. Is the necessary sample size impossible to get? Then reduce power or widen your minimum detectable effect. Negotiate with reality. Reality won’t budge; perhaps your design will.

In the end though, the MDE itself is an exercise in humility. It is an acknowledgement that there are things beyond your ability to spot. You need to decide what does and doesn’t matter. Otherwise, a test becomes nothing more than a coin flip with some overhead. It may yield something but you’ll always wonder if that something is worth anything.

Decide ahead of time what the thing is you’re looking for and then let the data do its work. The picture comes into focus. Everything else is simply waiting on the light.

Minimum Detectable Effect Calculator