Overview
This lecture follows on from the previous week’s session on randomised controlled trials and covers what systematic reviews and meta-analyses are, why doctors need them, and how they are conducted and appraised. It uses one Cochrane review as a worked example throughout - Taylor et al., “Statins for the primary prevention of cardiovascular disease” (Cochrane Database of Systematic Reviews 2013, Issue 1, CD004816) - to illustrate the seven steps of a systematic review, risk-of-bias and publication-bias assessment, meta-analysis and forest plots, subgroup/sensitivity analysis, when a meta-analysis should not be done, and reporting standards (PRISMA).
What is a systematic review, and when should it be used?
A systematic review “attempts to identify, appraise and synthesize all the empirical evidence that meets pre-specified eligibility criteria to answer a specific research question,” using explicit, systematic methods chosen to minimise bias, to produce more reliable findings for decision-making (Cochrane definition).
It includes:
- Identification of relevant studies from a number of different sources, including unpublished sources
- Selection of studies for inclusion, with evaluation of their strengths and limitations against clear, predefined criteria
- Systematic collection of data
- Appropriate synthesis of data
When a doctor might want to read one:
- To answer a clinical question
- If there are many studies relating to the clinical question
- If the results of studies seem to contradict each other
- To inform the planning of research, by identifying gaps in knowledge
- Because systematic reviews are vital for producing evidence-based guidelines
When a doctor should not rely on one:
- If it answers a narrow question that doesn’t relate to your clinical question
- If the authors haven’t followed best-practice guidelines for conducting systematic reviews
- If it doesn’t answer your clinical question in a way that would be acceptable to the patient
Not all systematic reviews are Cochrane reviews, and not all systematic reviews are of good quality; this is why guidelines exist covering registration, conduct and reporting.
Systematic review vs narrative review
| Feature | Narrative review | Systematic review |
|---|---|---|
| Question | Often broad in scope | Often a focused clinical question |
| Sources and search | Not usually specified, potentially biased | Comprehensive sources, explicit search strategy |
| Selection | Not usually specified, potentially biased | Criterion-based selection, uniformly applied |
| Appraisal | Variable | Rigorous critical appraisal |
| Synthesis | Often a qualitative summary | Quantitative summary* |
| Inferences | Sometimes evidence-based | Usually evidence-based |
*A quantitative summary that includes a statistical synthesis is a meta-analysis.
Narrative reviews may be heavily influenced by opinion; systematic reviews aim to be replicable, transparent and systematic. [transcript note: the lecture raises “what about scoping reviews??” as an open question and does not answer it]
The Cochrane Collaboration and quality safeguards
The Cochrane Collaboration is an international network headquartered in the UK, a registered not-for-profit and a member of the UK National Council for Voluntary Organizations, with members and supporters from more than 190 countries. It has gathered and summarised research evidence for 30 years and does not accept commercial or conflicted funding, which is described as vital to generating authoritative and reliable information, free from commercial and financial constraint.
Types of Cochrane review:
- Intervention reviews - effectiveness/safety of a treatment, vaccine, device, preventative measure, procedure or policy
- Diagnostic test accuracy reviews - accuracy of a test, device or scale to aid diagnosis
- Prognosis reviews - describe and predict the course of a disease/health condition
- Qualitative evidence syntheses - perspectives and experiences of an intervention or condition
- Methodology reviews - how research is designed, conducted, reported or used
- Overviews of reviews - synthesise multiple systematic reviews on related questions
- Rapid reviews - systematic reviews accelerated by streamlining or omitting specific methods
- Prototype reviews - other review types without established standard Cochrane methodology yet, e.g. scoping reviews, mixed-methods reviews, reviews of prevalence studies, realist reviews
The Cochrane Library hosts the Cochrane Database of Systematic Reviews, with search and browse functions, editorials and special collections. PROSPERO (run via the University of York) is an international database of prospectively registered systematic review protocols across health, social care, welfare, public health, education, crime, justice and international development, intended to reduce duplication of effort and reporting bias.
A systematic review (with or without meta-analysis) is itself a piece of research, needing a multidisciplinary team: typically a librarian, several reviewers, and a biostatistician. More than one person should extract data from each report, since independent double extraction produces fewer errors than single extraction followed by a second person’s verification (Buscemi et al 2006); one study found data extraction errors in 20 of 34 reviews (Jones et al 2005).
The seven steps of a systematic review: question, protocol and search
- Formulation of a clear question
- Write a protocol for the review and register it
- Search for relevant studies
- Collect data from studies
- Assessment of included studies
- Synthesis of findings
- Interpretation of results
Step 1, clear question: define precisely the question the review will address. In the worked example, reducing high blood cholesterol (a CVD risk factor) is an important pharmacotherapy goal; statins are first-choice agents. Earlier reviews showed benefit in people with existing CVD, but the case for primary prevention was uncertain when the previous review version was published (2011), prompting this update. Objective: to assess the effects, both harms and benefits, of statins in people with no history of CVD.
Step 2, write and register a protocol: a protocol states the review’s objectives and planned methods before the review is carried out (illustrated with a different example protocol, on proton pump inhibitors in preterm infants with GORD).
Step 3, search for relevant studies: requires pre-specified eligibility criteria. In the worked example:
- Types of studies: RCTs comparing statins for at least 12 months with placebo or usual care; outcome follow-up of at least 6 months
- Types of participants: men and women aged 18+, no restriction on cholesterol levels; ≤10% of the study population could have prior CVD (angina, MI and/or stroke); trials using statins to treat chronic conditions (e.g. Alzheimer’s disease, rheumatoid arthritis, renal disease, macular degeneration, aortic stenosis) were excluded
- Types of interventions: statins (HMG CoA reductase inhibitors) vs placebo or usual care
- Concomitant interventions: allowed if given to both arms equally; one additional adjuvant drug allowed if a patient developed excessively high lipids during the trial
- Outcome measures collected: death from all causes; fatal and non-fatal CHD, CVD and stroke events; combined endpoint; change in total and LDL cholesterol; revascularisation; adverse events; quality of life; costs [transcript note: the slide circles some of these outcomes and strikes through others, but does not explain what the annotation means]
A fully documented search strategy (shown from a different Cochrane review, on saturated fat and CVD, as an example of expected reporting detail) combines cholesterol/lipid-lowering search terms with diet/nutrition terms and cardiovascular-disease MeSH terms, with no language restrictions. [transcript note: the right-hand end of one search line is cut off at the slide edge]
Result of the search (PRISMA-style flow diagram, statin review update): 6,439 records identified through database searching (duplicates removed) plus 3 records from other sources gave 6,442 records screened; 6,311 were excluded; 131 full-text papers were retrieved; 39 were excluded, with reasons; 92 full-text articles were included, comprising 35 articles on 7 existing trials and 57 articles on 5 new trials, plus 1 trial awaiting classification, and 56 articles on 4 new trials.
Data extraction and risk-of-bias assessment
Step 4, data extraction: two reviewers independently extract and record data using a systematic approach and standard format; disagreements are resolved by a pre-specified process (e.g. discussion, or a third person arbitrating). Data extracted includes trial methods, participants, outcomes, outcome data, and risk of bias.
Worked example, PREVEND IT 2004: a randomised, 2x2 factorial trial of 864 participants with microalbuminuria in Holland, aged 28-75 (mean 51), 64.5% men, 96% Caucasian, less than 10% with clinical evidence of CVD. Intervention: 40 mg pravastatin vs placebo, followed up for 3.8 years. Primary outcome: composite of fatal and non-fatal CVD events; single outcomes included fatal CVD events, stroke, heart failure, non-fatal MI and cholesterol.
Step 5, assessment of included studies (risk of bias), judged for this study as:
- Random sequence generation: low risk (computer-generated randomisation)
- Allocation concealment: low risk (participants allocated to a treatment number)
- Blinding (performance and detection bias): low risk (double blind)
- Incomplete outcome data (attrition bias): unclear risk (intention-to-treat used but confined to CVD events; 6% dropped out)
[transcript note: this risk-of-bias table continues beyond what is shown, so further domains for this study are not recorded here]
Across all included studies, a methodological quality graph summarised, for six domains (random sequence generation, allocation concealment, blinding, incomplete outcome data, selective reporting, other bias), the percentage of studies judged low, unclear or high risk of bias; most domains were majority low/unclear risk, with “other bias” trending towards more unclear/high risk.
Publication bias and funnel plots
Reporting biases can arise when a review is based only on published studies, or only on studies published in English. Drug trials may go unpublished if they show no effect (publication bias), and many RCTs, especially with conflicts of interest, carry a high risk of bias. Because access to all relevant data (particularly drug trial data) is needed, reviewers search trial registers, such as the ANZCTR (Australian New Zealand Clinical Trials Registry), for unpublished trials.
A funnel plot is used to assess publication/reporting bias for a given outcome:
- X-axis: odds ratio (OR) on a logarithmic scale, so 0.5 and 2.0 are equidistant from 1.0
- Y-axis: standard error of log OR (smaller studies, with fewer outcomes, sit lower because they have larger SE)
- Dashed vertical line: pooled estimate of the treatment effect
- Solid line: line of no effect (null value)
A roughly symmetric, funnel-shaped scatter of study points suggests no strong reporting bias.
One example funnel plot in the lecture shows extra open-circle points alongside the filled circles, apparently representing imputed or missing studies used to illustrate asymmetry from publication bias, but this is not explicitly labelled on the slide.
What is meta-analysis, and when should it be done?
A meta-analysis is an optional part of a systematic review: a statistical analysis that combines results from two or more separate studies, calculating a weighted average of the effect estimates reported in the individual studies. More weight is given to studies that provide more information: more participants, more events, or lower variance (less uncertainty).
Why use meta-analysis:
- To increase power and precision: better ability to detect a true relationship if one exists, and narrower confidence intervals from a larger pooled sample size
- To assess consistency and generalisability of results, by quantifying between-study variation (heterogeneity)
- It may make it possible to answer questions the original studies didn’t pose (e.g. does the effect vary by age group, a subgroup analysis), to resolve controversies from conflicting studies, or to generate new hypotheses
When a meta-analysis can be done:
- When more than one study has estimated an effect
- When study characteristics are similar enough that combining them makes sense
- When data are available in a form that allows combination (e.g. outcomes measured in similar ways)
Meta-analysis can combine group-level or individual-participant-level data.
Reading forest plots
A forest plot shows each study’s effect estimate and confidence interval as a box and line (box size reflecting the study’s weight/precision), with a pooled estimate shown as a diamond at the bottom.
Worked example, “Total Number of CHD Events” (statin therapy vs usual care/placebo, 13 studies: ACAPS 1994, Adult Japanese MEGA Study, AFCAPS/TexCAPS 1998, ASPEN 2006, CAIUS 1996, CARDS 2008, CERDIA 2004, HYRIM 2007, JUPITER 2008, KAPS 1995, METEOR 2010, PHYLLIS 2004, PREVEND IT 2004, WOSCOPS): pooled risk ratio 0.73 (95% CI 0.67 to 0.80), 820/24,217 events (statin) vs 1114/23,832 (control). Heterogeneity: Chi² = 14.48, df = 13 (P = 0.34), I² = 10%. Test for overall effect: Z = 7.07 (P < 0.00001). The pooled diamond sits to the left of the line of no effect, favouring statin treatment.
Worked example, adverse events - myalgia/muscle pain (8 studies): risk ratios mostly straddled 1. Pooled risk ratio 1.03 (95% CI 0.97 to 1.09), 1847/19,396 events (statin) vs 1704/18,542 (control). Heterogeneity: Chi² = 13.53, df = 8 (P = 0.09), I² = 41%. Test for overall effect: Z = 0.84 (P = 0.40), i.e. no significant difference in myalgia/muscle pain risk between groups.
Subgroup analysis and sensitivity analysis
Subgroup analysis is used where it is suspected in advance that certain features may alter the effect of an intervention. Sensitivity analysis asks whether the result changes according to small variations in the data and the methods used. [transcript note: the slide introducing these two concepts has bullet text that appears truncated/incomplete as transcribed]
Worked example, subgroup analysis (individual-participant-data meta-analysis of statins in older people, the Lancet CTT paper introduced at the start of the lecture): effects on major vascular event components per 1 mmol/L reduction in LDL cholesterol, by age band at randomisation (≤55, >55-≤60, >60-≤65, >65-≤70, >70-<75, >75 years). For major coronary events, the trend test across age groups was significant (χ² = 6.92, p = 0.009); risk ratio in the oldest group (>75 years) was 0.82 (0.70-0.96), versus an overall pooled RR of 0.76 (0.73-0.79), i.e. benefit persisted in the oldest age group but appeared somewhat attenuated.
Worked example, subgroup analysis within a single RCT (SPRINT trial, “A Randomized Trial of Intensive versus Standard Blood-Pressure Control”): a forest plot of hazard ratios (intensive vs standard treatment) across subgroups (previous chronic kidney disease, age <75/≥75, sex, race, previous cardiovascular disease, systolic blood pressure band), each with an interaction p-value; none were statistically significant (approximately 0.32-0.83), indicating the intensive-treatment benefit was broadly consistent across subgroups.
When not to do a meta-analysis
Quality issues:
- A meta-analysis is only as good as the studies in it (“rubbish in, rubbish out”)
- Narrow confidence intervals around the combined result of biased studies can be worse than the biased studies on their own, by creating false confidence
- Meta-analysis in the presence of serious publication and/or reporting bias may produce an inappropriate summary
Studies too diverse:
- Combining studies that are too heterogeneous is not useful; the pooled result may be meaningless, and genuine effects in individual studies may be obscured
Interpreting results and reporting standards
Step 7, interpretation of results:
- Discussion: statement of principal findings; strengths and limitations of the review; interpretation of the review (strengths and limitations of the evidence; direction and magnitude of the overall effect estimate, e.g. relative risk or mean difference, from the summarised studies)
- Conclusions
- Recommendations, including implications for practice and research
Recommendations from a systematic review are not necessarily best for individual patients.
Good reporting is crucial. PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) is an evidence-based minimum set of items for reporting systematic reviews and meta-analyses, for use by authors and by peer reviewers/editors. Related resources include the PRISMA 2020 checklist and flow diagram, PROSPERO (for registration), and the EQUATOR Network.
As far as shown, the PRISMA 2020 checklist covers: Title (identify the report as a systematic review); Abstract; Introduction (rationale in the context of existing knowledge; explicit statement of objectives/questions); and Methods, including eligibility criteria and how studies were grouped for synthesis, information sources searched (with search dates), full search strategies, the study selection process (number of reviewers, independence, automation tools), the data collection process, data items sought (outcomes and other variables, and how missing/unclear information was handled), risk-of-bias assessment methods, effect measures used for each outcome, and synthesis methods. [transcript note: the checklist table shown is cut off after Methods item 13a; the remaining items, and the Results, Discussion and Other information sections, are not shown]
Self-test
- Define a systematic review, using Cochrane’s definition.
- List the four things a systematic review includes.
- Describe three situations in which a doctor might want to read a systematic review.
- Describe three situations in which a doctor should be cautious about relying on a systematic review.
- Distinguish a systematic review from a narrative review, across question, sources/search, selection, appraisal and synthesis.
- What feature of Cochrane’s funding is highlighted as important, and why?
- List four types of Cochrane review and what each assesses.
- What is PROSPERO, and what problem does registering a review’s protocol there help reduce?
- Why is independent double data extraction recommended over single extraction with verification?
- List, in order, the seven steps of conducting a systematic review.
- In the worked statin review, describe the eligibility criteria for types of studies, participants and interventions.
- What outcome measures were collected in the worked example’s search?
- Using the worked example’s numbers, describe what a PRISMA-style study flow diagram shows, from records identified through to studies included.
- Describe the risk-of-bias domains assessed for the PREVEND IT 2004 study, and the judgement given for each.
- Explain how to read a funnel plot: what do the axes represent, and what do the dashed and solid lines represent?
- A funnel plot shows several small studies apparently missing from one side, with no matching studies on the other side. What might this suggest, and why?
- Define meta-analysis, and describe how a study’s “weight” within it is determined.
- Describe two reasons for performing a meta-analysis, beyond simply combining data.
- List the three conditions needed before a meta-analysis can be done.
- In a forest plot, what does the diamond represent, and what does it mean if it sits entirely to the left of the line of no effect?
- Using the “Total Number of CHD Events” forest plot, state the pooled risk ratio (with 95% CI) and explain what the I² value indicated about heterogeneity.
- Distinguish subgroup analysis from sensitivity analysis.
- In the SPRINT subgroup forest plot, what did the non-significant interaction p-values indicate about the intensive-treatment effect?
- List two reasons a meta-analysis should not be performed.
- Describe the three components of Step 7 (interpretation of results) in a systematic review.
- What is PRISMA, and why does the lecture describe good reporting as crucial?
- A clinician finds a systematic review whose funnel plot looks strongly asymmetric, but whose forest plot shows a large, statistically significant pooled effect with low heterogeneity (I² = 5%). Should the low heterogeneity reassure the clinician? Explain.
Answers
Reveal answers
- A systematic review attempts to identify, appraise and synthesize all the empirical evidence that meets pre-specified eligibility criteria to answer a specific research question, using explicit, systematic methods selected to minimise bias, to produce more reliable findings for decision-making.
- Identification of relevant studies from multiple sources, including unpublished ones; selection of studies for inclusion with evaluation of strengths/limitations against clear, predefined criteria; systematic collection of data; appropriate synthesis of data.
- Any three of: to answer a clinical question; when many studies relate to the clinical question; when study results seem to contradict each other; to inform research planning by identifying knowledge gaps; because they are vital for producing evidence-based guidelines.
- Any three of: it answers a narrow question unrelated to your clinical question; the authors haven’t followed best-practice guidelines for conducting systematic reviews; it doesn’t answer your clinical question in a way acceptable to the patient.
- Narrative reviews: often a broad question, unspecified/potentially biased sources and selection, variable appraisal, often a qualitative synthesis. Systematic reviews: a focused clinical question, comprehensive sources with an explicit search strategy, criterion-based selection uniformly applied, rigorous critical appraisal, and a quantitative summary (a quantitative synthesis with statistical pooling is a meta-analysis).
- Cochrane does not accept commercial or conflicted funding; this independence is described as vital to generating authoritative, reliable information, free of commercial and financial constraint.
- Any four of: intervention reviews (effectiveness/safety of a treatment, vaccine, device, preventive measure, procedure or policy); diagnostic test accuracy reviews (accuracy of a test/device/scale for diagnosis); prognosis reviews (describe/predict disease course); qualitative evidence syntheses (perspectives/experiences of an intervention or condition); methodology reviews (how research is designed/conducted/reported/used); overviews of reviews (synthesise multiple systematic reviews); rapid reviews (accelerated by streamlining/omitting methods); prototype reviews (types without established Cochrane methodology yet, e.g. scoping reviews).
- PROSPERO is an international database of prospectively registered systematic review protocols, across health, social care, welfare, public health, education, crime, justice and international development; registering there is intended to reduce duplication of effort and reporting bias.
- Independent double extraction produces fewer errors than single extraction followed by a second person’s verification (Buscemi et al 2006), and data extraction errors have been found to be common, in 20 of 34 reviews in one study (Jones et al 2005).
- Formulation of a clear question; write a protocol for the review and register it; search for relevant studies; collect data from studies; assessment of included studies; synthesis of findings; interpretation of results.
- Types of studies: RCTs comparing statins for at least 12 months with placebo/usual care, with at least 6 months’ outcome follow-up. Types of participants: adults aged 18+, no cholesterol-level restriction, ≤10% prior CVD history, excluding trials treating chronic conditions like Alzheimer’s disease, rheumatoid arthritis, renal disease, macular degeneration or aortic stenosis. Types of interventions: statins (HMG CoA reductase inhibitors) vs placebo or usual care.
- Death from all causes; fatal and non-fatal CHD, CVD and stroke events; combined endpoint; change in total and LDL cholesterol; revascularisation; adverse events; quality of life; costs.
- 6,439 records from database searching (duplicates removed) plus 3 from other sources gave 6,442 records screened; 6,311 were excluded; 131 full-text papers were retrieved; 39 were excluded with reasons; 92 full-text articles were included, comprising 35 articles on 7 existing trials and 57 articles on 5 new trials, plus 1 trial awaiting classification, and 56 articles on 4 new trials.
- Random sequence generation: low risk (computer-generated randomisation). Allocation concealment: low risk (participants allocated to a treatment number). Blinding: low risk (double blind). Incomplete outcome data: unclear risk (intention-to-treat used but confined to CVD events, with 6% drop-out).
- X-axis: odds ratio on a log scale, so 0.5 and 2.0 sit equidistant from 1.0. Y-axis: standard error of log OR, with smaller (more uncertain) studies lower down. Dashed vertical line: the pooled treatment-effect estimate. Solid line: the line of no effect (null value).
- It suggests possible publication or reporting bias: smaller studies showing no effect, or an unfavourable effect, may not have been published, producing asymmetry in the funnel around the pooled estimate.
- Meta-analysis is a statistical analysis combining results from two or more separate studies, calculating a weighted average of their effect estimates. More weight goes to studies with more participants, more events, or lower variance (less uncertainty).
- Any two of: increasing power and precision (better ability to detect a true effect, narrower CIs from a larger pooled sample); assessing consistency/generalisability by quantifying heterogeneity; answering questions the original studies didn’t pose (e.g. subgroup analysis by age); resolving controversy from conflicting studies; generating new hypotheses.
- More than one study has estimated an effect; the studies’ characteristics are similar enough to combine sensibly; the data are available and in a combinable form (e.g. outcomes measured similarly).
- The diamond represents the pooled treatment-effect estimate across all studies; sitting entirely to the left of the line of no effect means the pooled result favours treatment over control.
- Pooled risk ratio 0.73 (95% CI 0.67 to 0.80). I² = 10%, indicating low heterogeneity between the 13 studies.
- Subgroup analysis examines whether an intervention’s effect differs where certain features are suspected in advance to alter it (e.g. by age); sensitivity analysis examines whether the result changes with small variations in the data or methods used.
- None of the interaction p-values were statistically significant, indicating the intensive blood-pressure treatment’s benefit was broadly consistent across the subgroups examined.
- Any two of: the included studies are of poor quality (narrow CIs around a biased pooled result can be worse than the biased studies alone, “rubbish in, rubbish out”); there is serious publication and/or reporting bias; the included studies are too diverse to usefully combine.
- Discussion (statement of principal findings; strengths and limitations of the review; interpretation of the review, including strengths/limitations of the evidence and the direction/magnitude of the overall effect estimate); conclusions; recommendations, including implications for practice and research, noting these are not necessarily best for individual patients.
- PRISMA is an evidence-based minimum set of items for reporting systematic reviews and meta-analyses, for authors and for peer reviewers/editors to check against. Good reporting is crucial because it is what allows a review’s methods and findings to be assessed, replicated and trusted.
- No, not necessarily. Low I² only shows the individual study estimates were numerically consistent with one another; it says nothing about whether unpublished or negative studies were missing. A strongly asymmetric funnel plot is itself a warning sign of publication/reporting bias, and pooling a biased set of studies can produce a falsely narrow, falsely confident combined estimate (“rubbish in, rubbish out”) even when the included studies agree closely with each other.