Conducting a Systematic Review: Selection, Data Collection, and Risk of Bias

Systematic Review: Searching for Studies

  • The goal of the searching phase in a systematic review is to identify and synthesize all relevant high-quality evidence that answers the specific research question, regardless of whether that evidence has been published or not.

  • A comprehensive search strategy is mandatory and includes:

    • Utilizing electronic databases to locate relevant research publications.

    • Employing various other methods to identify both published and unpublished studies.

  • A highly sensitive and well-constructed search strategy includes:

    • Multiple content terms using MeSH (Medical Subject Headings) and text word filters.

    • A sensitive methods filter (e.g., specific filters for RCTs).

    • The appropriate use of Boolean operators (AND, OR, NOT).

  • Reproducibility is paramount; the search strategy must be described in enough detail to allow others to replicate it.

  • Searching should extend beyond standard databases to include supplementary literature, often called "grey literature," which includes:

    • Trial registries, such as the WHO International Clinical Trials Registry platform, clinicaltrials.gov, and the Australia New Zealand Clinical Trial Registry.

    • Reference lists of retrieved articles.

    • Reference lists of previous relevant systematic reviews.

    • Scientific meeting and conference abstracts.

    • Citation index searching.

    • Contact with experts in the field.

    • Contact with relevant industry or pharmaceutical companies.

    • Dissertations and theses databases.

  • Example of a search strategy description:

    • The authors searched Central, Medline, Embase, and Cenar.

    • No language restrictions were imposed.

    • They reviewed previous systematic reviews, editorials, and narrative reviews for relevant references.

    • They searched trial registries and contacted specialists for unpublished data.

    • The full search strategy was provided in an appendix for transparency.

Selecting Studies for Inclusion

  • Study selection must be based on pre-specified inclusion or eligibility criteria explicitly stated in the systematic review protocol.

  • The selection process should be conducted by at least 22 people acting independently.

  • Disagreements regarding study inclusion are typically resolved through discussion between the two reviewers, with a 3rd3^{rd} reviewer available for arbitration if necessary.

  • Study flow diagrams (often PRISMA-style) are used to summarize the search results and the selection process.

  • The unit of analysis for reviews of interventions is the individual Randomized Controlled Trial (RCT). Reviewers must take care to avoid duplication, as one study may result in multiple publications.

  • Example Case Study of the Selection Process:

    • Studies carried forward from a previous review: 6363

    • Records identified through database searching: 13,00013,000

    • Records identified through trial registries: 300300

    • Total records screened (by title and abstract after duplicate removal): 11,00011,000

    • Full-text articles assessed for eligibility: 244244

    • Full-text articles excluded: 201201 (with explicit reasons provided).

    • New studies included: 2222 (across 4343 separate publications).

    • Total studies in qualitative analysis: 8585

    • Total studies in quantitative analysis: 7575

Data Collection Procedures

  • Data collection is one of the most critical and time-intensive phases of a systematic review.

  • To minimize errors and bias, more than one person should extract data independently, especially when the data being collected is subjective.

  • A standardized data collection form should be used to ensure the systematic and consistent gathering of information.

  • Information extracted from individual RCTs includes:

    • Study design and conduct.

    • Individual study results.

    • PICO M categories: Methods (study design, type), Participants, Intervention, Control, and Outcomes (both primary and secondary outcomes).

  • Example of data extraction protocol:

    • 22 authors independently extracted characteristics using a standardized form.

    • A 3rd3^{rd} author checked all extracted data for accuracy.

    • Consensus was reached for any disagreements.

Assessing Risk of Bias in Included Studies

  • A valid systematic review is dependent on the validity of the underlying studies. If included studies are biased, the meta-analysis may produce precise but misleading summary estimates.

  • Every systematic review must include a Risk of Bias (RoB) assessment for all included studies.

  • Assessment should be conducted by at least 22 independent reviewers, with a 3rd3^{rd} reviewer for arbitration of disagreements.

  • Domain-based assessments, specifically the Cochrane risk of bias tools, are recommended to evaluate potential sources of bias.

  • Original Cochrane Risk of Bias Domains:

    • Random sequence generation.

    • Allocation concealment.

    • Blinding of participants and personnel.

    • Blinding of outcome assessors.

    • Incomplete outcome data.

    • Selective outcome reporting.

    • Other potential sources of bias.

  • Risk of Bias 22 (released in 20192019) was designed to reduce jargon and clarify exactly where and why bias occurs in RCT reporting and conduct.

    • Domain 11: Bias arising from the randomization process (Was the sequence random? Was allocation concealed until assignment? Were there suspicious baseline differences?).

    • Domain 22: Bias due to deviations from intended interventions (Awareness of participants/personnel, appropriateness of analysis to estimate effect of assignment).

    • Domain 33: Bias due to missing outcome data (Data available for all or nearly all randomized participants).

    • Domain 44: Bias in measurement of the outcome (Appropriateness of measurement method, differences in ascertainment between groups).

    • Domain 55: Bias in selection of the reported result (Analysis in accordance with a pre-specified plan finalized before unblinding).

  • Visualization of Risk of Bias:

    • Risk of Bias Summary: A table listing studies by rows and domains by columns; colored icons indicate Low Risk (green), Some Concerns (yellow), or High Risk (red).

    • Risk of Bias Graph: Represents judgments as percentages across all included studies (mostly green signifies high quality; mostly red signifies low quality).

  • Outcome-Specific Assessment: Assessments may vary by outcome. For example, a study on phototherapy for atopic eczema might have different risk profiles for "physician-assessed clinical signs" versus "patient-reported symptoms."

  • Overall Judgment Levels in RoB 22:

    • Low risk of bias: Low for all domains.

    • Some concerns: At least 11 domain raises concern, but none are high risk.

    • High risk of bias: At least 11 domain is high risk, or multiple domains have concerns that substantially lower confidence.

GRADE: Grading Quality of Evidence

  • GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) provides guidance for rating the quality of evidence and the strength of healthcare recommendations.

  • GRADE applies overall quality ratings per outcome across a body of evidence.

  • Quality Levels:

    • High: High confidence that the true effect is similar to the estimated effect.

    • Moderate: The true effect is probably close to the estimated effect.

    • Low: The true effect might be markedly different from the estimated effect.

    • Very Low: The true effect is probably markedly different from the estimated effect.

  • Baseline Standards: RCT evidence starts at "High" quality; observational data starts at "Low" quality due to residual confounding.

  • Factors for Rating Down Certainty:

    • High risk of bias in the base studies.

    • Imprecision (e.g., small sample sizes, few events, or wide confidence intervals).

    • Inconsistency/Heterogeneity (variation in results across studies).

    • Indirectness (studies do not directly address the review question).

    • Publication bias (suspicion of selective publication of results).

  • Factors for Rating Up Certainty:

    • Large magnitude of effect.

    • Presence of a dose-response gradient.

    • Results that oppose plausible residual confounding.

  • GRADE is reported in the text and in the "Summary of Findings" table.

  • Example (Exercise-based cardiac rehabilitation):

    • All-cause mortality: Moderate certainty evidence.

    • Cardiovascular mortality: Moderate certainty evidence.

    • Fatal and/or non-fatal myocardial infarction: High certainty evidence.