Overview

This lecture follows on from the previous week’s session on randomised controlled trials and covers what systematic reviews and meta-analyses are, why doctors need them, and how they are conducted and appraised. It uses one Cochrane review as a worked example throughout - Taylor et al., “Statins for the primary prevention of cardiovascular disease” (Cochrane Database of Systematic Reviews 2013, Issue 1, CD004816) - to illustrate the seven steps of a systematic review, risk-of-bias and publication-bias assessment, meta-analysis and forest plots, subgroup/sensitivity analysis, when a meta-analysis should not be done, and reporting standards (PRISMA).

What is a systematic review, and when should it be used?

A systematic review “attempts to identify, appraise and synthesize all the empirical evidence that meets pre-specified eligibility criteria to answer a specific research question,” using explicit, systematic methods chosen to minimise bias, to produce more reliable findings for decision-making (Cochrane definition).

It includes:

  • Identification of relevant studies from a number of different sources, including unpublished sources
  • Selection of studies for inclusion, with evaluation of their strengths and limitations against clear, predefined criteria
  • Systematic collection of data
  • Appropriate synthesis of data

When a doctor might want to read one:

  • To answer a clinical question
  • If there are many studies relating to the clinical question
  • If the results of studies seem to contradict each other
  • To inform the planning of research, by identifying gaps in knowledge
  • Because systematic reviews are vital for producing evidence-based guidelines

When a doctor should not rely on one:

  • If it answers a narrow question that doesn’t relate to your clinical question
  • If the authors haven’t followed best-practice guidelines for conducting systematic reviews
  • If it doesn’t answer your clinical question in a way that would be acceptable to the patient

Not all systematic reviews are Cochrane reviews, and not all systematic reviews are of good quality; this is why guidelines exist covering registration, conduct and reporting.

Systematic review vs narrative review

FeatureNarrative reviewSystematic review
QuestionOften broad in scopeOften a focused clinical question
Sources and searchNot usually specified, potentially biasedComprehensive sources, explicit search strategy
SelectionNot usually specified, potentially biasedCriterion-based selection, uniformly applied
AppraisalVariableRigorous critical appraisal
SynthesisOften a qualitative summaryQuantitative summary*
InferencesSometimes evidence-basedUsually evidence-based

*A quantitative summary that includes a statistical synthesis is a meta-analysis.

Narrative reviews may be heavily influenced by opinion; systematic reviews aim to be replicable, transparent and systematic. [transcript note: the lecture raises “what about scoping reviews??” as an open question and does not answer it]

The Cochrane Collaboration and quality safeguards

The Cochrane Collaboration is an international network headquartered in the UK, a registered not-for-profit and a member of the UK National Council for Voluntary Organizations, with members and supporters from more than 190 countries. It has gathered and summarised research evidence for 30 years and does not accept commercial or conflicted funding, which is described as vital to generating authoritative and reliable information, free from commercial and financial constraint.

Types of Cochrane review:

  • Intervention reviews - effectiveness/safety of a treatment, vaccine, device, preventative measure, procedure or policy
  • Diagnostic test accuracy reviews - accuracy of a test, device or scale to aid diagnosis
  • Prognosis reviews - describe and predict the course of a disease/health condition
  • Qualitative evidence syntheses - perspectives and experiences of an intervention or condition
  • Methodology reviews - how research is designed, conducted, reported or used
  • Overviews of reviews - synthesise multiple systematic reviews on related questions
  • Rapid reviews - systematic reviews accelerated by streamlining or omitting specific methods
  • Prototype reviews - other review types without established standard Cochrane methodology yet, e.g. scoping reviews, mixed-methods reviews, reviews of prevalence studies, realist reviews

The Cochrane Library hosts the Cochrane Database of Systematic Reviews, with search and browse functions, editorials and special collections. PROSPERO (run via the University of York) is an international database of prospectively registered systematic review protocols across health, social care, welfare, public health, education, crime, justice and international development, intended to reduce duplication of effort and reporting bias.

A systematic review (with or without meta-analysis) is itself a piece of research, needing a multidisciplinary team: typically a librarian, several reviewers, and a biostatistician. More than one person should extract data from each report, since independent double extraction produces fewer errors than single extraction followed by a second person’s verification (Buscemi et al 2006); one study found data extraction errors in 20 of 34 reviews (Jones et al 2005).

  1. Formulation of a clear question
  2. Write a protocol for the review and register it
  3. Search for relevant studies
  4. Collect data from studies
  5. Assessment of included studies
  6. Synthesis of findings
  7. Interpretation of results

Step 1, clear question: define precisely the question the review will address. In the worked example, reducing high blood cholesterol (a CVD risk factor) is an important pharmacotherapy goal; statins are first-choice agents. Earlier reviews showed benefit in people with existing CVD, but the case for primary prevention was uncertain when the previous review version was published (2011), prompting this update. Objective: to assess the effects, both harms and benefits, of statins in people with no history of CVD.

Step 2, write and register a protocol: a protocol states the review’s objectives and planned methods before the review is carried out (illustrated with a different example protocol, on proton pump inhibitors in preterm infants with GORD).

Step 3, search for relevant studies: requires pre-specified eligibility criteria. In the worked example:

  • Types of studies: RCTs comparing statins for at least 12 months with placebo or usual care; outcome follow-up of at least 6 months
  • Types of participants: men and women aged 18+, no restriction on cholesterol levels; ≤10% of the study population could have prior CVD (angina, MI and/or stroke); trials using statins to treat chronic conditions (e.g. Alzheimer’s disease, rheumatoid arthritis, renal disease, macular degeneration, aortic stenosis) were excluded
  • Types of interventions: statins (HMG CoA reductase inhibitors) vs placebo or usual care
  • Concomitant interventions: allowed if given to both arms equally; one additional adjuvant drug allowed if a patient developed excessively high lipids during the trial
  • Outcome measures collected: death from all causes; fatal and non-fatal CHD, CVD and stroke events; combined endpoint; change in total and LDL cholesterol; revascularisation; adverse events; quality of life; costs [transcript note: the slide circles some of these outcomes and strikes through others, but does not explain what the annotation means]

A fully documented search strategy (shown from a different Cochrane review, on saturated fat and CVD, as an example of expected reporting detail) combines cholesterol/lipid-lowering search terms with diet/nutrition terms and cardiovascular-disease MeSH terms, with no language restrictions. [transcript note: the right-hand end of one search line is cut off at the slide edge]

Result of the search (PRISMA-style flow diagram, statin review update): 6,439 records identified through database searching (duplicates removed) plus 3 records from other sources gave 6,442 records screened; 6,311 were excluded; 131 full-text papers were retrieved; 39 were excluded, with reasons; 92 full-text articles were included, comprising 35 articles on 7 existing trials and 57 articles on 5 new trials, plus 1 trial awaiting classification, and 56 articles on 4 new trials.

Data extraction and risk-of-bias assessment

Step 4, data extraction: two reviewers independently extract and record data using a systematic approach and standard format; disagreements are resolved by a pre-specified process (e.g. discussion, or a third person arbitrating). Data extracted includes trial methods, participants, outcomes, outcome data, and risk of bias.

Worked example, PREVEND IT 2004: a randomised, 2x2 factorial trial of 864 participants with microalbuminuria in Holland, aged 28-75 (mean 51), 64.5% men, 96% Caucasian, less than 10% with clinical evidence of CVD. Intervention: 40 mg pravastatin vs placebo, followed up for 3.8 years. Primary outcome: composite of fatal and non-fatal CVD events; single outcomes included fatal CVD events, stroke, heart failure, non-fatal MI and cholesterol.

Step 5, assessment of included studies (risk of bias), judged for this study as:

  • Random sequence generation: low risk (computer-generated randomisation)
  • Allocation concealment: low risk (participants allocated to a treatment number)
  • Blinding (performance and detection bias): low risk (double blind)
  • Incomplete outcome data (attrition bias): unclear risk (intention-to-treat used but confined to CVD events; 6% dropped out)

[transcript note: this risk-of-bias table continues beyond what is shown, so further domains for this study are not recorded here]

Across all included studies, a methodological quality graph summarised, for six domains (random sequence generation, allocation concealment, blinding, incomplete outcome data, selective reporting, other bias), the percentage of studies judged low, unclear or high risk of bias; most domains were majority low/unclear risk, with “other bias” trending towards more unclear/high risk.

Publication bias and funnel plots

Reporting biases can arise when a review is based only on published studies, or only on studies published in English. Drug trials may go unpublished if they show no effect (publication bias), and many RCTs, especially with conflicts of interest, carry a high risk of bias. Because access to all relevant data (particularly drug trial data) is needed, reviewers search trial registers, such as the ANZCTR (Australian New Zealand Clinical Trials Registry), for unpublished trials.

A funnel plot is used to assess publication/reporting bias for a given outcome:

  • X-axis: odds ratio (OR) on a logarithmic scale, so 0.5 and 2.0 are equidistant from 1.0
  • Y-axis: standard error of log OR (smaller studies, with fewer outcomes, sit lower because they have larger SE)
  • Dashed vertical line: pooled estimate of the treatment effect
  • Solid line: line of no effect (null value)

A roughly symmetric, funnel-shaped scatter of study points suggests no strong reporting bias.

One example funnel plot in the lecture shows extra open-circle points alongside the filled circles, apparently representing imputed or missing studies used to illustrate asymmetry from publication bias, but this is not explicitly labelled on the slide.

What is meta-analysis, and when should it be done?

A meta-analysis is an optional part of a systematic review: a statistical analysis that combines results from two or more separate studies, calculating a weighted average of the effect estimates reported in the individual studies. More weight is given to studies that provide more information: more participants, more events, or lower variance (less uncertainty).

Why use meta-analysis:

  • To increase power and precision: better ability to detect a true relationship if one exists, and narrower confidence intervals from a larger pooled sample size
  • To assess consistency and generalisability of results, by quantifying between-study variation (heterogeneity)
  • It may make it possible to answer questions the original studies didn’t pose (e.g. does the effect vary by age group, a subgroup analysis), to resolve controversies from conflicting studies, or to generate new hypotheses

When a meta-analysis can be done:

  • When more than one study has estimated an effect
  • When study characteristics are similar enough that combining them makes sense
  • When data are available in a form that allows combination (e.g. outcomes measured in similar ways)

Meta-analysis can combine group-level or individual-participant-level data.

Reading forest plots

A forest plot shows each study’s effect estimate and confidence interval as a box and line (box size reflecting the study’s weight/precision), with a pooled estimate shown as a diamond at the bottom.

Worked example, “Total Number of CHD Events” (statin therapy vs usual care/placebo, 13 studies: ACAPS 1994, Adult Japanese MEGA Study, AFCAPS/TexCAPS 1998, ASPEN 2006, CAIUS 1996, CARDS 2008, CERDIA 2004, HYRIM 2007, JUPITER 2008, KAPS 1995, METEOR 2010, PHYLLIS 2004, PREVEND IT 2004, WOSCOPS): pooled risk ratio 0.73 (95% CI 0.67 to 0.80), 820/24,217 events (statin) vs 1114/23,832 (control). Heterogeneity: Chi² = 14.48, df = 13 (P = 0.34), I² = 10%. Test for overall effect: Z = 7.07 (P < 0.00001). The pooled diamond sits to the left of the line of no effect, favouring statin treatment.

Worked example, adverse events - myalgia/muscle pain (8 studies): risk ratios mostly straddled 1. Pooled risk ratio 1.03 (95% CI 0.97 to 1.09), 1847/19,396 events (statin) vs 1704/18,542 (control). Heterogeneity: Chi² = 13.53, df = 8 (P = 0.09), I² = 41%. Test for overall effect: Z = 0.84 (P = 0.40), i.e. no significant difference in myalgia/muscle pain risk between groups.

Subgroup analysis and sensitivity analysis

Subgroup analysis is used where it is suspected in advance that certain features may alter the effect of an intervention. Sensitivity analysis asks whether the result changes according to small variations in the data and the methods used. [transcript note: the slide introducing these two concepts has bullet text that appears truncated/incomplete as transcribed]

Worked example, subgroup analysis (individual-participant-data meta-analysis of statins in older people, the Lancet CTT paper introduced at the start of the lecture): effects on major vascular event components per 1 mmol/L reduction in LDL cholesterol, by age band at randomisation (≤55, >55-≤60, >60-≤65, >65-≤70, >70-<75, >75 years). For major coronary events, the trend test across age groups was significant (χ² = 6.92, p = 0.009); risk ratio in the oldest group (>75 years) was 0.82 (0.70-0.96), versus an overall pooled RR of 0.76 (0.73-0.79), i.e. benefit persisted in the oldest age group but appeared somewhat attenuated.

Worked example, subgroup analysis within a single RCT (SPRINT trial, “A Randomized Trial of Intensive versus Standard Blood-Pressure Control”): a forest plot of hazard ratios (intensive vs standard treatment) across subgroups (previous chronic kidney disease, age <75/≥75, sex, race, previous cardiovascular disease, systolic blood pressure band), each with an interaction p-value; none were statistically significant (approximately 0.32-0.83), indicating the intensive-treatment benefit was broadly consistent across subgroups.

When not to do a meta-analysis

Quality issues:

  • A meta-analysis is only as good as the studies in it (“rubbish in, rubbish out”)
  • Narrow confidence intervals around the combined result of biased studies can be worse than the biased studies on their own, by creating false confidence
  • Meta-analysis in the presence of serious publication and/or reporting bias may produce an inappropriate summary

Studies too diverse:

  • Combining studies that are too heterogeneous is not useful; the pooled result may be meaningless, and genuine effects in individual studies may be obscured

Interpreting results and reporting standards

Step 7, interpretation of results:

  • Discussion: statement of principal findings; strengths and limitations of the review; interpretation of the review (strengths and limitations of the evidence; direction and magnitude of the overall effect estimate, e.g. relative risk or mean difference, from the summarised studies)
  • Conclusions
  • Recommendations, including implications for practice and research

Recommendations from a systematic review are not necessarily best for individual patients.

Good reporting is crucial. PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) is an evidence-based minimum set of items for reporting systematic reviews and meta-analyses, for use by authors and by peer reviewers/editors. Related resources include the PRISMA 2020 checklist and flow diagram, PROSPERO (for registration), and the EQUATOR Network.

As far as shown, the PRISMA 2020 checklist covers: Title (identify the report as a systematic review); Abstract; Introduction (rationale in the context of existing knowledge; explicit statement of objectives/questions); and Methods, including eligibility criteria and how studies were grouped for synthesis, information sources searched (with search dates), full search strategies, the study selection process (number of reviewers, independence, automation tools), the data collection process, data items sought (outcomes and other variables, and how missing/unclear information was handled), risk-of-bias assessment methods, effect measures used for each outcome, and synthesis methods. [transcript note: the checklist table shown is cut off after Methods item 13a; the remaining items, and the Results, Discussion and Other information sections, are not shown]

Self-test

  1. Define a systematic review, using Cochrane’s definition.
  2. List the four things a systematic review includes.
  3. Describe three situations in which a doctor might want to read a systematic review.
  4. Describe three situations in which a doctor should be cautious about relying on a systematic review.
  5. Distinguish a systematic review from a narrative review, across question, sources/search, selection, appraisal and synthesis.
  6. What feature of Cochrane’s funding is highlighted as important, and why?
  7. List four types of Cochrane review and what each assesses.
  8. What is PROSPERO, and what problem does registering a review’s protocol there help reduce?
  9. Why is independent double data extraction recommended over single extraction with verification?
  10. List, in order, the seven steps of conducting a systematic review.
  11. In the worked statin review, describe the eligibility criteria for types of studies, participants and interventions.
  12. What outcome measures were collected in the worked example’s search?
  13. Using the worked example’s numbers, describe what a PRISMA-style study flow diagram shows, from records identified through to studies included.
  14. Describe the risk-of-bias domains assessed for the PREVEND IT 2004 study, and the judgement given for each.
  15. Explain how to read a funnel plot: what do the axes represent, and what do the dashed and solid lines represent?
  16. A funnel plot shows several small studies apparently missing from one side, with no matching studies on the other side. What might this suggest, and why?
  17. Define meta-analysis, and describe how a study’s “weight” within it is determined.
  18. Describe two reasons for performing a meta-analysis, beyond simply combining data.
  19. List the three conditions needed before a meta-analysis can be done.
  20. In a forest plot, what does the diamond represent, and what does it mean if it sits entirely to the left of the line of no effect?
  21. Using the “Total Number of CHD Events” forest plot, state the pooled risk ratio (with 95% CI) and explain what the I² value indicated about heterogeneity.
  22. Distinguish subgroup analysis from sensitivity analysis.
  23. In the SPRINT subgroup forest plot, what did the non-significant interaction p-values indicate about the intensive-treatment effect?
  24. List two reasons a meta-analysis should not be performed.
  25. Describe the three components of Step 7 (interpretation of results) in a systematic review.
  26. What is PRISMA, and why does the lecture describe good reporting as crucial?
  27. A clinician finds a systematic review whose funnel plot looks strongly asymmetric, but whose forest plot shows a large, statistically significant pooled effect with low heterogeneity (I² = 5%). Should the low heterogeneity reassure the clinician? Explain.

Answers