Systematic Reviews
Dual Screening Without Bias: A Systematic Review Workflow That Survives Peer Review
Published: 2026-01-12
A practical guide to title/abstract and full-text screening in systematic reviews — independent dual screening, resolving disagreements, and reporting it well.
Dual Screening Without Bias: A Systematic Review Workflow That Survives Peer Review
Peer reviewers read the methods section of a systematic review before they read anything else, and screening is usually the first thing they scrutinize. A single line — "one reviewer screened all records" — is enough to trigger a major revision request or an outright rejection at some journals, no matter how solid the rest of the analysis is. Screening isn't a clerical step before the "real" work starts; it's where selection bias either gets controlled or quietly built into the review.
This article walks through a screening workflow that holds up under scrutiny: how to set it up before you start, how to run title/abstract and full-text screening properly, how to resolve disagreements in a way that's defensible, and how to report it so a reviewer doesn't have to ask.
Why Independent, Dual Screening Matters
The PRISMA 2020 statement — the current reporting standard for systematic reviews, published in BMJ — reflects a broad methodological consensus that study selection should be done independently by at least two reviewers, with clear reporting of how many reviewers screened records and how disagreements were resolved. The reasoning is straightforward: a single reviewer's judgment about whether a borderline abstract meets your inclusion criteria is inevitably shaped by their own reading of the topic, their familiarity with certain authors or methods, and simple fatigue during a long screening session. Two independent reviewers, screening without seeing each other's decisions first, catch each other's errors and inconsistencies.
The evidence backs this up directly. A 2019 methodological systematic review by Waffenschmidt and colleagues, published in BMC Medical Research Methodology, compared single- against double-screening across published evaluations and found that single screening missed a median of 5% of eligible studies, with the range running as high as 58% depending on reviewer experience — and that less experienced reviewers working alone missed substantially more than experienced ones. In several cases, the missed studies were enough to have changed the pooled result. Single screening isn't equivalent to dual screening; it's a different, riskier method that needs to be justified and disclosed if used at all.
Setting Up the Workflow Before You Start Screening
The single biggest predictor of a smooth screening process is time spent before the first record is screened.
Write eligibility criteria that don't require judgment calls. "Studies of moderate to severe disease" is an invitation for two reviewers to disagree; "studies enrolling patients with a documented diagnosis of X using criteria Y" is not. The more your PICO-based eligibility criteria read like inclusion/exclusion rules rather than descriptions, the less arbitration you'll need later.
Run a calibration round. Before screening the full set, have both reviewers independently screen the same sample of 50–100 records, then meet to compare decisions and discuss discrepancies. This step routinely surfaces ambiguity in the eligibility criteria while it's still cheap to fix — before it has propagated across thousands of records.
Pick a screening tool that tracks independence. Platforms such as Rayyan, Covidence, DistillerSR, or EPPI-Reviewer allow two reviewers to screen blinded to each other's decisions and then reveal conflicts automatically. A shared spreadsheet can work for a small review, but it makes true blinding harder to enforce and audit later — reviewers can see each other's marks, consciously or not.
Stage 1: Title and Abstract Screening
Both reviewers screen every record independently against the pre-defined eligibility criteria, applying a simple decision at this stage: include, exclude, or unsure. Because abstracts often don't contain enough detail to apply every criterion, the "unsure" category exists on purpose — it should be used generously rather than forcing a premature exclusion. Anything marked "include" or "unsure" by either reviewer proceeds to full-text review; only records both reviewers independently excluded are dropped at this stage.
Detailed, record-by-record reasons for exclusion are not required at title/abstract stage under PRISMA 2020 — that level of documentation is expected at full-text stage — but keeping a rough sense of the most common reasons is useful for troubleshooting the criteria as you go.
Stage 2: Full-Text Screening
The same independent, dual-reviewer principle applies, but now every excluded study needs a documented, specific reason — "wrong population," "wrong comparator," "outcome not reported," "duplicate publication," and so on. This reason list becomes the basis for the exclusion counts reported in the PRISMA flow diagram, which readers use to understand exactly how the final study count was reached. Vague reasons ("did not meet inclusion criteria") are a common cause of reviewer pushback, because they give readers no way to judge whether the exclusion was appropriate.
Resolving Disagreements: Building an Audit Trail
Disagreements between reviewers are expected and are not a sign of poor screening — a screening process with zero disagreements more often signals that reviewers weren't truly independent than that the criteria were unusually clear. What matters is how disagreements are resolved and documented.
Discussion first. Most conflicts resolve through a short conversation once both reviewers explain their reasoning against the eligibility criteria — often surfacing a criterion that needs to be clarified for the rest of the screening.
A third reviewer for anything unresolved. When two reviewers can't agree, a pre-specified third team member — ideally someone senior enough to arbitrate, who wasn't involved in the original screening decision — makes the final call. Decide who this will be, and how they'll be briefed, before screening begins, not after the first disagreement appears.
Quantify agreement. Reporting Cohen's kappa for title/abstract screening agreement gives readers and reviewers a concrete sense of screening reliability, beyond just stating that "two reviewers screened independently." It also gives you an early warning sign: a low kappa partway through screening is a signal to pause and recalibrate, not to push through and hope it improves.
Reporting Screening Methodology for Peer Review
When a manuscript reaches peer review, reviewers with systematic review experience will specifically look for:
- The number of reviewers who screened titles/abstracts and full texts, and whether they worked independently
- How disagreements were resolved (discussion, third reviewer, or both) and by whom
- A completed PRISMA 2020 flow diagram showing records identified, duplicates removed, records screened, records excluded (with reasons at full-text stage), and studies included
- Whether any automation or single-reviewer shortcuts were used, and — if so — an explicit justification, since PRISMA 2020 expects transparency about deviations from dual independent screening rather than silent omission
Writing this section in the methods as a short, specific paragraph — reviewer count, independence, tool used, conflict resolution process, and a pointer to the flow diagram — resolves most reviewer questions before they're asked.
Common Screening Mistakes That Trigger Revision Requests
- Reporting "two reviewers screened" without stating independence. If reviewers screened together or could see each other's decisions in real time, that's not independent dual screening — and reviewers with methodological expertise will usually ask directly whether it was.
- No documented process for resolving disagreements. Silence here reads as either an oversight or an unreported single-reviewer override.
- Full-text exclusions without individual reasons, making the flow diagram numbers impossible to audit.
- Changing eligibility criteria mid-screening without noting it, which undermines the reproducibility that dual screening is meant to protect.
- Treating the calibration round as optional. Skipping it tends to surface as inconsistent early decisions that have to be revisited later, costing more time than the calibration round would have.
A transparent, well-documented screening process is one of the few parts of a systematic review that a peer reviewer can verify almost entirely from the methods section and flow diagram alone — which is exactly why it's worth getting right the first time.
Frequently Asked Questions
Do both reviewers need to screen 100% of records?
Yes, for conventional dual screening as reflected in PRISMA 2020 guidance. Single-reviewer screening, or "single screening with verification" of only a sample, is a recognized but riskier alternative that should be explicitly justified and reported as a limitation, not presented as equivalent to full dual screening.
What is Cohen's kappa and what counts as a good score?
Cohen's kappa is a statistic that measures agreement between two reviewers beyond what would be expected by chance alone, ranging from -1 to 1. Values above roughly 0.6 are generally considered substantial agreement in systematic review screening, though the acceptable threshold can vary by field and should be interpreted alongside the actual disagreements, not in isolation.
What tools are commonly used for systematic review screening?
Rayyan, Covidence, DistillerSR, and EPPI-Reviewer are widely used platforms that support blinded, independent dual screening and automatically flag conflicts for resolution. The right choice often depends on institutional access and review size rather than one tool being universally best.
Do I need a third reviewer for every systematic review?
You need a pre-specified plan for resolving disagreements that discussion alone doesn't settle. A third reviewer is the most common approach, but the important part is deciding who that will be, and their role, before screening starts.
How does PRISMA 2020 address screening specifically?
PRISMA 2020 asks authors to report the screening process in enough detail for it to be reproduced, including the number of reviewers, whether they worked independently, how disagreements were resolved, and a flow diagram accounting for every record from identification through final inclusion.
Is it ever acceptable for one reviewer to screen and another to verify only a sample?
It can be, particularly for time-constrained reviews, but the evidence shows it is not equivalent to full dual screening and carries a real risk of missed studies. If used, it should be explicitly reported as a methodological choice, ideally with a sensitivity check on the sample that was verified.
References
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.
- Waffenschmidt S, Knelangen M, Sieben W, Bühn S, Pieper D. Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Medical Research Methodology. 2019;19:132.
- Higgins JPT, Thomas J, Chandler J, et al. (editors). Cochrane Handbook for Systematic Reviews of Interventions. Cochrane.
- Higgins JPT, Green S (editors, earlier editions); Cochrane training materials on study selection and screening tools.
