How do I calculate the sample size for a study?
Decide the smallest effect that would matter, the significance level, and the power you want, then run a power calculation for your planned analysis. The hard part is the effect size, since it comes from prior research or from a judgment about what would be practically meaningful, not from the software.
Updated
The calculation itself takes minutes in any statistics package. Everything difficult is upstream of it, in the numbers you feed in, and that is where the reviewer questions land.
The effect size is the hard input. Three sources are defensible: a meta-analysis or prior study in a comparable population, a pilot study of your own, or a judgment about the smallest effect that would be practically meaningful. The third is often the best, because it ties the study to a decision rather than to a previous estimate that may itself be inflated.
Be careful using a published effect size directly. Published effects are systematically larger than true effects, because studies that found large effects are more likely to be published, so powering a study on a published estimate frequently produces a study that is underpowered for the real effect. Assuming something smaller is the safer choice.
How much that input matters is easy to underestimate. Holding power at 0.80, Serdar and colleagues report that a standardized effect of 0.20 requires 788 observations, while an effect of 2.5 reaches roughly 0.80 power with 8 (Serdar et al., 2021, p. 9). Halving your assumed effect does not add a few participants to the total, it multiplies it. The number you defend to a reviewer is the effect size, and the sample follows from it.
Then add for attrition and exclusions. A calculation returning 128 participants means 128 in the analysis, not 128 recruited. If you expect 20 percent dropout, recruit 160, and say so in the protocol.
If your study is a descriptive survey rather than a hypothesis test, the inputs change to population size, margin of error, and confidence level. For a population of 10,000, a 95 percent confidence level with a 5 percent margin of error gives 370 respondents. Move the confidence level to 90 or 99 percent and the requirement becomes 264 or 623. Hold confidence at 95 percent and change the margin of error instead, and it becomes 96 at 10 percent or 4,900 at 1 percent (Serdar et al., 2021, p. 22). Precision, not confidence, is what drives the recruitment cost.
None of this transfers to interview-based work, where the equivalent question is when to stop. Vasileiou and colleagues report saturation at interview 12 in one homogeneous and tightly focused study, 20 to 40 interviews where meta-themes had to hold across sites, and 17 interviews where the theoretical constructs were specified in advance. The same review separates code saturation at around 9 interviews from meaning saturation at 16 to 24 (Vasileiou et al., 2018, p. 3). Those are not rival answers to one question, they track scope, heterogeneity, number of sites, and whether you want to name issues or understand them in depth. Saturation accounted for 55 percent of the sample size justifications in that review, and none of those claims was substantiated through the study procedures the authors reported (Vasileiou et al., 2018, p. 15). If you cite saturation, show the analytic steps that produced it.
What is a power analysis actually doing?
Working out how many observations you need for your test to detect an effect of a stated size, given a chosen false-positive rate, with a stated probability. Power of 80 percent means that if the effect is really that size, you would find it in eight studies out of ten.
Four quantities are linked: sample size, effect size, significance level, and power. Fix any three and the fourth follows, which is why the same tool can answer either "how many do I need" or "what could I have detected".
This is also why a conventional number like 30 fails as a default. In a two-group comparison with Gaussian outcomes, alpha at 0.05 and power at 0.80, a standardized effect of 1 does return a total of 30 (Serdar et al., 2021, p. 4). Fix the sample at 30 and shrink the effect to 0.2, and the same design is badly underpowered (Serdar et al., 2021, p. 9). The 30 was never a rule, it was the answer to one particular set of assumptions.
The version to avoid is post hoc power, calculating power from the effect you observed after the study. It is a restatement of your p value rather than new information, and reviewers who know this will say so.
What do I do if I cannot reach the number?
Say so and adjust the claim. Report the effect size your achievable sample can detect, and frame the study around that. A small study reported honestly, with confidence intervals and no significance claims it cannot support, is publishable. One that hides its power is not.
There are also design routes that recover power without more participants: repeated measures within the same people, more reliable measurement instruments, a more homogeneous sample, or a continuous outcome instead of a dichotomized one. Dichotomizing a continuous measure is a surprisingly common and expensive way to lose power.
Where the study is genuinely underpowered and must go ahead, treat it as estimation rather than testing. Report the estimate and its interval, describe the study as preliminary, and let the interval speak about the uncertainty.
Sources
- Serdar et al., Sample size, power and effect size revisited: simplified and practical approaches in pre-clinical, clinical and laboratory studies, Biochemia Medica (2021) · checked 6 August 2026
- Greenland et al., Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations, European Journal of Epidemiology (2016) · checked 6 August 2026
- Vasileiou et al., Characterising and justifying sample size sufficiency in interview-based studies: systematic analysis of qualitative health research over a 15-year period, BMC Medical Research Methodology (2018) · checked 6 August 2026