Systematic Reviews

How to Handle Missing Data in a Systematic Review

July 11, 2026·Dr. Samuel Osei·5 min read
On this page

Almost every systematic review, regardless of topic, eventually runs into an included study that reports an outcome incompletely -- a mean without a standard deviation, a subgroup result without the overall sample size, or an outcome mentioned in a study's methods section but never actually reported in its results. How you handle these gaps is a real methodological decision that deserves the same transparency as your search strategy or risk-of-bias assessment, not a quiet workaround buried in a footnote.

The first step: contact the study authors

Before applying any statistical workaround, the most defensible first step is directly contacting the study's authors requesting the missing data. This is standard, expected practice in rigorous systematic reviews, and PRISMA 2020 specifically asks you to report how many studies you contacted and how many responded with usable data. Document this process explicitly -- the number of authors contacted, the response rate, and how missing data was handled for studies where authors did not respond or no longer had access to the original data.

Distinguishing types of missing data

Missing data at the outcome level (a study reports some outcomes but not the one you need) is different from missing data at the summary-statistic level (a study reports a result but not the standard deviation or confidence interval needed to include it in a meta-analysis), and both differ from missing participant-level data due to loss to follow-up within a study itself. Each type calls for a different approach, and conflating them in your methods section -- treating all missing data as if it were the same problem -- reads as imprecise methodology to a careful reviewer.

Imputing missing standard deviations

When a study reports a mean but not a standard deviation, several accepted imputation methods exist rather than simply excluding the study. Using the average standard deviation from other included studies measuring the same outcome, converting a reported standard error or confidence interval into a standard deviation using standard formulas, or using a reported p-value together with the sample size to back-calculate an approximate standard deviation are all recognized approaches, each with different assumptions and appropriateness depending on what information the study does report. Whichever method you use, state it explicitly in your methods section and apply it consistently across all studies with the same type of gap, rather than choosing a different workaround study by study.

Handling loss to follow-up within a trial

Participant dropout within an individual trial is a different problem from a study failing to report a summary statistic, and it should be addressed at the level of that trial's own risk-of-bias assessment rather than through pooling-stage imputation. A high or differential dropout rate between study arms is itself a risk-of-bias concern under both RoB 2 and ROBINS-I, and should be reflected in that study's individual risk-of-bias rating rather than treated purely as a statistical missing-data problem to be imputed around.

Sensitivity analysis: the honest safety net

Whatever imputation or handling method you choose for missing summary statistics, running a sensitivity analysis that excludes the affected studies entirely, and comparing the resulting pooled estimate to your primary analysis, is the clearest way to show readers how much your missing-data handling actually influenced your conclusion. If the pooled estimate barely shifts when the imputed studies are excluded, this is reassuring evidence your handling method did not meaningfully distort the result. If it shifts substantially, this is important information to report explicitly rather than downplay, since it tells readers your conclusion is sensitive to a methodological choice rather than robust across it.

What GRADE expects regarding missing data

Missing data at a level serious enough to raise real doubt about a pooled estimate's reliability is one of the considerations that can contribute to a GRADE imprecision or risk-of-bias downgrade, particularly when missingness is differential between comparison groups (suggesting outcome-related dropout) rather than random. Your GRADE assessment and your missing-data handling section should connect explicitly -- if you downgraded certainty partly due to missing data concerns, your results section should make clear exactly which studies and which specific gaps drove that judgment.

What not to do

Simply excluding studies with incomplete outcome reporting from your pooled analysis without disclosure, or silently substituting a value without stating your method and rationale, are both practices that undermine reproducibility and are exactly the kind of undisclosed methodological choice a careful peer reviewer is trained to look for. Transparent, consistent, and disclosed handling of missing data -- even an imperfect method, honestly reported -- is considerably more defensible than an undisclosed gap discovered later by someone trying to reproduce your analysis.

Missing studies versus missing data within studies

It's worth distinguishing missing data within an otherwise-included study from an entire study your search should have found but did not -- the latter is a search comprehensiveness or publication bias issue, addressed through your search strategy and funnel plot analysis, not through the data-handling methods discussed here. Conflating these two genuinely different problems in your methods section, treating a comprehensiveness gap as if it were simply a missing-statistic problem to impute around, misrepresents what your specific methodology can and genuinely cannot address, and a careful, experienced reviewer will generally notice fairly quickly when a genuine search-comprehensiveness issue has been quietly relabeled and folded into an imputation footnote instead of being addressed directly and honestly on its own terms.

Documenting your approach in the PRISMA flow diagram and text

Where missing data affected your ability to include a study's outcome in meta-analysis specifically, this should be traceable through your PRISMA flow diagram or an accompanying table, showing which studies were formally eligible but excluded from the quantitative synthesis specifically due to unresolvable missing data, distinct from studies excluded earlier at the screening stage for simply not meeting your eligibility criteria in the first place, which is a genuinely different and separate methodological safeguard from the missing-data handling covered here, and conflating the two in your writeup weakens both explanations unnecessarily.

#missing data#systematic reviews#meta-analysis