Systematic Reviews

Dual Screening in Systematic Reviews: Why and How to Do It Properly

July 20, 2026·Dr. Lauren Ito·5 min read
On this page

Dual, independent screening -- two reviewers each assessing every record for eligibility without seeing the other's decisions until a comparison stage -- is one of the methodological choices most directly scrutinized by peer reviewers, because it speaks directly to whether your included studies were selected consistently and without individual bias.

Why single-reviewer screening is a real problem

A single reviewer screening thousands of titles and abstracts will make inconsistent judgment calls over the course of the process, simply due to fatigue, evolving familiarity with the literature, and the natural drift in how strictly eligibility criteria get applied across a long screening session. There is also no way to distinguish a reviewer's genuine, defensible judgment call from a simple error without a second, independent assessment to compare against.

Cochrane Handbook guidance and most major reporting standards now treat dual independent screening as an expected default for at least the title and abstract stage, and increasingly for full-text screening as well, particularly for reviews intended for high-impact publication or reviews informing clinical or policy decisions.

How the process actually works

Two reviewers independently screen the same set of records against the same pre-specified eligibility criteria, without discussing individual records or seeing each other's decisions during the initial pass. This independence is the entire point -- if reviewers discuss records as they go, you no longer have two genuinely independent assessments, you have one assessment with a running consensus discussion, which does not let you measure or report actual inter-rater agreement.

After both reviewers complete their independent screening, decisions are compared. Records where both reviewers agree on inclusion or exclusion proceed accordingly. Records with disagreement go to a resolution step, commonly involving direct discussion between the two reviewers to reach consensus, or escalation to a third reviewer who was not involved in the initial screening to make a tie-breaking decision.

Measuring and reporting agreement

Cohen's kappa is the standard statistic for quantifying inter-rater agreement between two screeners, accounting for the level of agreement expected by chance alone rather than simply reporting a raw percentage agreement, which can look artificially high or low depending on how many records fall into each decision category. Kappa values are typically interpreted using established bands, from poor or slight agreement at the low end to almost perfect agreement at the high end, though as with I-squared thresholds in meta-analysis, these bands are conventions rather than hard rules.

Reporting your kappa statistic in the methods section is expected practice and signals to reviewers that you actually measured screening consistency rather than simply asserting that dual screening occurred. A conspicuously absent kappa statistic, when dual screening is claimed, is a gap methods reviewers specifically look for.

Full-text screening deserves the same rigor

Title and abstract screening filters a large volume of records quickly using limited information, and disagreements at this stage are common and expected, since abstracts often do not contain enough detail to make a fully confident eligibility judgment. Full-text screening, applied to a much smaller shortlist, should still be conducted independently by two reviewers where feasible, since this stage is where eligibility criteria get applied with full information, and errors here directly determine your final included study list.

Documenting conflict resolution in your methods

Your methods section should specify not just that dual screening occurred, but exactly how disagreements were resolved -- discussion to consensus, third-reviewer adjudication, or a pre-specified decision rule. This should be decided and documented in your protocol before screening begins, for the same reason your eligibility criteria are pre-specified: to prevent the resolution process itself from becoming a source of undisclosed, inconsistent judgment calls.

What this looks like in your PRISMA flow diagram

Your PRISMA flow diagram should reflect the outcome of the dual-screening and resolution process at each stage -- records excluded at title and abstract, records assessed at full-text, and the specific reasons full-text exclusions occurred. A flow diagram whose numbers do not obviously connect to a documented dual-screening and resolution process is one of the most common places reviewers find inconsistencies, because the numbers at each stage should be fully traceable back to your described screening methodology.

Practical constraints and honest reporting

Dual independent screening genuinely takes more reviewer time than single-reviewer screening, and resource constraints are real, particularly for student researchers or small teams. If genuine dual screening of the entire record set is not feasible, a common and defensible compromise is dual screening of a representative percentage of records -- commonly 10 to 20 percent -- to establish and report an agreement rate, with the remainder screened by a single reviewer. This should be disclosed explicitly and honestly as a limitation, rather than presented as if full dual screening occurred when it did not,

Screening software built for this workflow

Dedicated systematic review screening platforms, including Covidence and Rayyan among others, are built specifically to support dual independent screening -- each reviewer screens blind to the other's decisions within the platform, and the software automatically flags disagreements for resolution rather than requiring a manual comparison of two separate spreadsheets. Beyond convenience, using purpose-built software creates a cleaner audit trail of exactly when each decision was made and by whom, which is useful both for your own methods reporting and for responding to a reviewer question about your screening process months after the work was actually done.

Dual independent screening genuinely takes more reviewer time than single-reviewer screening, and resource constraints are real, particularly for student researchers or small teams. If genuine dual screening of the entire record set is not feasible, a common and defensible compromise is dual screening of a representative percentage of records -- commonly 10 to 20 percent -- to establish and report an agreement rate, with the remainder screened by a single reviewer. This should be disclosed explicitly and honestly as a limitation, rather than presented as if full dual screening occurred when it did not, since an undisclosed gap between claimed and actual methodology is a far more serious credibility problem than an honestly reported resource constraint.

#screening#systematic reviews#PRISMA