AI in Medical Diagnosis: Building a Defensible Systematic Review Protocol
On this page
- Framing the question in diagnostic accuracy terms
- Why QUADAS-2 is the starting point, not RoB 2
- When PROBAST applies instead
- Extraction fields specific to diagnostic accuracy
- Statistical synthesis for diagnostic accuracy data
- Reference standard variability as a specific quality concern
- Search strategy spanning clinical and technical literature
- Reporting with PRISMA-DTA
- A practical starting checklist
- Handling studies with incomplete accuracy reporting
- Considering clinical utility beyond raw accuracy figures
- A closing consideration on regulatory relevance
- A final word on communicating uncertainty honestly
- Considering the clinicians who will actually read your findings
Systematic reviews examining AI's role in medical diagnosis are, at their methodological core, diagnostic test accuracy reviews -- evaluating how well an AI tool identifies a condition compared to an established reference standard -- and recognizing this from the protocol stage shapes nearly every subsequent decision differently than a standard intervention-effectiveness review would.
Framing the question in diagnostic accuracy terms
Rather than a standard PICO structure, a diagnostic accuracy review question specifies the AI tool or model being evaluated, the target condition it aims to identify, the reference standard it is being compared against, and the specific population in which this comparison is being made. A vaguely framed question like "how accurate is AI at diagnosis" needs this specific structure before a genuinely searchable, screenable systematic review question exists.
Why QUADAS-2 is the starting point, not RoB 2
Because these are fundamentally diagnostic accuracy studies, QUADAS-2, not RoB 2, is the appropriate risk-of-bias tool for the primary studies typically included in this kind of review, addressing patient selection, the index test's conduct, the reference standard's adequacy, and flow and timing through the study -- concerns specific to diagnostic accuracy research that RoB 2's intervention-trial-focused domains do not address.
When PROBAST applies instead
If your specific review question concerns an AI model's development and validation as a prediction tool, rather than its diagnostic accuracy against a reference standard in a defined clinical population, PROBAST is the more appropriate tool, and clarifying which of these two genuinely different research questions your review addresses is a necessary early decision, since the two tools assess different concerns and are not interchangeable.
Extraction fields specific to diagnostic accuracy
Your data extraction form needs fields capturing true positives, false positives, true negatives, and false negatives, or equivalently reported sensitivity and specificity with the underlying sample sizes, since this is the data structure diagnostic accuracy meta-analysis requires, genuinely different from the means, standard deviations, or event rates a standard intervention review would extract.
Statistical synthesis for diagnostic accuracy data
Where meta-analysis is appropriate, pooling sensitivity and specificity jointly, using a bivariate or hierarchical summary receiver operating characteristic model, is the standard approach, accounting for the known trade-off between these two measures rather than pooling them as though they were independent outcomes, a genuinely different statistical approach than standard risk-ratio or mean-difference meta-analysis.
Reference standard variability as a specific quality concern
A recurring quality issue in this specific literature is variability, or outright inadequacy, in what reference standard studies use to establish ground truth -- some studies compare the AI tool against expert clinician judgment, others against a more rigorous gold-standard diagnostic test, and this difference meaningfully affects how the reported accuracy figures should be interpreted and compared across your included studies.
Search strategy spanning clinical and technical literature
As with AI in healthcare more broadly, diagnostic AI research is published across both clinical journals and technical machine learning venues, and a comprehensive search needs to span both, including relevant technical databases and preprint servers alongside standard clinical databases like MEDLINE and Embase.
Reporting with PRISMA-DTA
The diagnostic test accuracy extension of PRISMA provides reporting guidance specifically suited to this review type, and following it explicitly, rather than standard PRISMA 2020 alone, signals to reviewers familiar with diagnostic accuracy methodology that your reporting genuinely reflects this review's actual structure.
A practical starting checklist
Before beginning your search, confirm you have framed your question using diagnostic accuracy structure rather than standard PICO, selected QUADAS-2 or PROBAST based on which specific question you are actually asking, and planned an extraction form capturing the true positive, false positive, true negative, and false negative data your intended synthesis method will need.
Handling studies with incomplete accuracy reporting
A recurring practical challenge in this literature is primary studies reporting only summary accuracy statistics like sensitivity and specificity without the underlying raw counts needed for certain meta-analytic approaches. Contacting original study authors for this underlying data, where feasible, or working with the summary statistics using appropriately adapted statistical methods, are both legitimate paths forward, and deciding which approach your specific review will take is worth addressing explicitly in your protocol rather than encountering this gap unprepared during analysis.
Considering clinical utility beyond raw accuracy figures
A tool with impressive standalone accuracy figures may still offer limited genuine clinical utility if it does not meaningfully change clinical decision-making or patient outcomes in practice, and where your included studies report this kind of downstream clinical impact data, alongside standalone accuracy figures, presenting both gives readers a fuller, more clinically meaningful picture than accuracy statistics in isolation.
A closing consideration on regulatory relevance
Given how directly diagnostic AI tools increasingly intersect with regulatory approval processes in many jurisdictions, a rigorously conducted systematic review on this topic can carry genuine weight beyond academic interest alone, informing real regulatory and clinical adoption decisions. Treating this responsibility with genuine seriousness is what separates a review that actually shapes practice from one that simply adds to the literature.
A final word on communicating uncertainty honestly
Diagnostic AI research often produces impressively precise-sounding accuracy figures, and resisting the temptation to present these with more confidence than your underlying evidence genuinely supports, particularly where included studies show meaningful heterogeneity or used less rigorous reference standards, is an essential discipline for this specific review type.
Considering the clinicians who will actually read your findings
A busy clinician deciding whether to trust a diagnostic AI tool needs a clear, honestly caveated answer more than an exhaustive technical discussion of statistical methods, and writing your discussion section with this practical reader specifically in mind, translating your findings into genuinely actionable guidance, is what makes a technically rigorous review also a genuinely useful one. That combination of rigor and readability is ultimately what separates a review that shapes real clinical decisions from one that simply adds to a growing pile of technical literature, one careful, honestly reported finding at a time, built for the clinicians who will actually rely on it.