Conducting a Systematic Review: Selection, Data Collection, and Risk of Bias
Systematic Review: Searching for Studies
The goal of the searching phase in a systematic review is to identify and synthesize all relevant high-quality evidence that answers the specific research question, regardless of whether that evidence has been published or not.
A comprehensive search strategy is mandatory and includes:
Utilizing electronic databases to locate relevant research publications.
Employing various other methods to identify both published and unpublished studies.
A highly sensitive and well-constructed search strategy includes:
Multiple content terms using MeSH (Medical Subject Headings) and text word filters.
A sensitive methods filter (e.g., specific filters for RCTs).
The appropriate use of Boolean operators (AND, OR, NOT).
Reproducibility is paramount; the search strategy must be described in enough detail to allow others to replicate it.
Searching should extend beyond standard databases to include supplementary literature, often called "grey literature," which includes:
Trial registries, such as the WHO International Clinical Trials Registry platform, clinicaltrials.gov, and the Australia New Zealand Clinical Trial Registry.
Reference lists of retrieved articles.
Reference lists of previous relevant systematic reviews.
Scientific meeting and conference abstracts.
Citation index searching.
Contact with experts in the field.
Contact with relevant industry or pharmaceutical companies.
Dissertations and theses databases.
Example of a search strategy description:
The authors searched Central, Medline, Embase, and Cenar.
No language restrictions were imposed.
They reviewed previous systematic reviews, editorials, and narrative reviews for relevant references.
They searched trial registries and contacted specialists for unpublished data.
The full search strategy was provided in an appendix for transparency.
Selecting Studies for Inclusion
Study selection must be based on pre-specified inclusion or eligibility criteria explicitly stated in the systematic review protocol.
The selection process should be conducted by at least people acting independently.
Disagreements regarding study inclusion are typically resolved through discussion between the two reviewers, with a reviewer available for arbitration if necessary.
Study flow diagrams (often PRISMA-style) are used to summarize the search results and the selection process.
The unit of analysis for reviews of interventions is the individual Randomized Controlled Trial (RCT). Reviewers must take care to avoid duplication, as one study may result in multiple publications.
Example Case Study of the Selection Process:
Studies carried forward from a previous review:
Records identified through database searching:
Records identified through trial registries:
Total records screened (by title and abstract after duplicate removal):
Full-text articles assessed for eligibility:
Full-text articles excluded: (with explicit reasons provided).
New studies included: (across separate publications).
Total studies in qualitative analysis:
Total studies in quantitative analysis:
Data Collection Procedures
Data collection is one of the most critical and time-intensive phases of a systematic review.
To minimize errors and bias, more than one person should extract data independently, especially when the data being collected is subjective.
A standardized data collection form should be used to ensure the systematic and consistent gathering of information.
Information extracted from individual RCTs includes:
Study design and conduct.
Individual study results.
PICO M categories: Methods (study design, type), Participants, Intervention, Control, and Outcomes (both primary and secondary outcomes).
Example of data extraction protocol:
authors independently extracted characteristics using a standardized form.
A author checked all extracted data for accuracy.
Consensus was reached for any disagreements.
Assessing Risk of Bias in Included Studies
A valid systematic review is dependent on the validity of the underlying studies. If included studies are biased, the meta-analysis may produce precise but misleading summary estimates.
Every systematic review must include a Risk of Bias (RoB) assessment for all included studies.
Assessment should be conducted by at least independent reviewers, with a reviewer for arbitration of disagreements.
Domain-based assessments, specifically the Cochrane risk of bias tools, are recommended to evaluate potential sources of bias.
Original Cochrane Risk of Bias Domains:
Random sequence generation.
Allocation concealment.
Blinding of participants and personnel.
Blinding of outcome assessors.
Incomplete outcome data.
Selective outcome reporting.
Other potential sources of bias.
Risk of Bias (released in ) was designed to reduce jargon and clarify exactly where and why bias occurs in RCT reporting and conduct.
Domain : Bias arising from the randomization process (Was the sequence random? Was allocation concealed until assignment? Were there suspicious baseline differences?).
Domain : Bias due to deviations from intended interventions (Awareness of participants/personnel, appropriateness of analysis to estimate effect of assignment).
Domain : Bias due to missing outcome data (Data available for all or nearly all randomized participants).
Domain : Bias in measurement of the outcome (Appropriateness of measurement method, differences in ascertainment between groups).
Domain : Bias in selection of the reported result (Analysis in accordance with a pre-specified plan finalized before unblinding).
Visualization of Risk of Bias:
Risk of Bias Summary: A table listing studies by rows and domains by columns; colored icons indicate Low Risk (green), Some Concerns (yellow), or High Risk (red).
Risk of Bias Graph: Represents judgments as percentages across all included studies (mostly green signifies high quality; mostly red signifies low quality).
Outcome-Specific Assessment: Assessments may vary by outcome. For example, a study on phototherapy for atopic eczema might have different risk profiles for "physician-assessed clinical signs" versus "patient-reported symptoms."
Overall Judgment Levels in RoB :
Low risk of bias: Low for all domains.
Some concerns: At least domain raises concern, but none are high risk.
High risk of bias: At least domain is high risk, or multiple domains have concerns that substantially lower confidence.
GRADE: Grading Quality of Evidence
GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) provides guidance for rating the quality of evidence and the strength of healthcare recommendations.
GRADE applies overall quality ratings per outcome across a body of evidence.
Quality Levels:
High: High confidence that the true effect is similar to the estimated effect.
Moderate: The true effect is probably close to the estimated effect.
Low: The true effect might be markedly different from the estimated effect.
Very Low: The true effect is probably markedly different from the estimated effect.
Baseline Standards: RCT evidence starts at "High" quality; observational data starts at "Low" quality due to residual confounding.
Factors for Rating Down Certainty:
High risk of bias in the base studies.
Imprecision (e.g., small sample sizes, few events, or wide confidence intervals).
Inconsistency/Heterogeneity (variation in results across studies).
Indirectness (studies do not directly address the review question).
Publication bias (suspicion of selective publication of results).
Factors for Rating Up Certainty:
Large magnitude of effect.
Presence of a dose-response gradient.
Results that oppose plausible residual confounding.
GRADE is reported in the text and in the "Summary of Findings" table.
Example (Exercise-based cardiac rehabilitation):
All-cause mortality: Moderate certainty evidence.
Cardiovascular mortality: Moderate certainty evidence.
Fatal and/or non-fatal myocardial infarction: High certainty evidence.