Overview

This lecture uses a clinical scenario (a 76 year old man whose statin was stopped in hospital) to frame two linked bodies of knowledge: how a randomised controlled trial should be designed, conducted and analysed, and how to read and critically appraise a published paper reporting one. The RCT half works through the phases of treatment evaluation, the anatomy of a trial, the ethical and regulatory context, choice of outcome measures and surrogate endpoints, duration of follow-up, and then a systematic treatment of bias (design, conduct, analysis and reporting) including randomisation, concealment of allocation, blinding, adherence, intention to treat analysis and the multiplicity problem. The appraisal half gives the anatomy of a paper section by section, the GATE picture, the source/eligible/participating population hierarchy, and the occurrence and effect estimates appropriate to categorical versus numerical outcomes. The lecture then returns to the scenario, appraises two statin trials in older adults, and resolves the decision with the evidence-based practice triad. The closing message is that an RCT is the best design for testing whether a treatment is beneficial, but trials can be done badly, so the results of a study cannot be accepted just because it is an RCT.

Clinical scenario: Mr X

  • It is 2024 and you are a GP. Mr X is a new patient: 76 year old man of European ethnicity, ex-smoker, father had a stroke aged 48 years.
  • Part of the assessment is cardiovascular risk assessment. He has no history of cardiovascular disease events (MI or stroke) but is concerned about his family history.
  • Past history: earlier in the year he was admitted to hospital with a head injury (passenger in a car hit by a truck). The hospital registrar stopped his statin but continued his blood pressure lowering medication. He is fully recovered and has come in for a prescription for his usual medications.
  • The two questions posed: why was the statin stopped, and should you re-start it?
  • On the recap, Mr X reports no side-effects, and that the registrar said something about his age.
  • Later social and values information: he likes to keep busy and has a small lawn-mowing business, he is a life-long tramper with an upcoming 3-day trip with the local tramping club, and he does not want to have a stroke like his father.

Warning

The transcript flags the recap slide (slide 82) as referencing case details introduced earlier in the lecture, outside the transcribed page range, and so not fully reconstructable from that slide alone.

CVD risk assessment and management in primary care

Ministry of Health (2018), Cardiovascular Disease Risk Assessment and Management for Primary Care. Lifestyle advice (diet, weight management, physical activity, smoking cessation) applies at every risk level. Note that the CVD risk prediction equations are validated for people under 75 years.

Cardiovascular riskDrug therapyFollow-up
Established CVDStrong evidence supports pharmacotherapy for modifiable risk factors, and antiplatelet therapy for secondary preventionReview annually
>15% CVD riskStrong evidence supports using statins and blood pressure lowering to prevent CVD events and deathsReview annually, repeat risk assessment annually
5–15% CVD riskDiscuss the magnitude of the benefits of statins or blood pressure lowering with the patient, based on the evidence that the higher the risk for the patient the more likely they are to benefitRisk 5–9%, repeat at five years. Risk 10–14%, repeat at two years
<5% CVD riskEvidence indicates medication management has limited benefitRisk <3%, repeat at 10 years. Risk 3–5%, repeat at five years

For people with diabetes an annual review is recommended.

Phases of treatment evaluation

  • RCTs sit at the apex of the evidence pyramid.
  • The development timeline runs pre-clinical → phase I → phase II (phase IIa and IIb) → phase III → phase IV, with milestones “first in humans”, phase I/II, phase II/III, and regulatory approval.
  • Phases I and II are a screening process to identify promising treatments and screen out treatments that do not look promising because of insufficient promise of efficacy or an unacceptably bad adverse event profile.
  • Phase III trials provide the definitive assessment of the effects of the intervention, on which licensing is based, and aim to determine whether there are tangible benefits to patients sufficient to offset any risks.
  • An evaluation of the intervention must be ethically acceptable, provide a reliable answer to a relevant clinical question, and be as efficient as possible.
  • Phase III trials are always randomised controlled trials.

Anatomy of a randomised controlled trial

The trial flow, in order:

  1. Patient population.
  2. Eligibility criteria and informed consent.
  3. Participants.
  4. Random allocation, splitting into experimental treatment and control.
  5. Follow up in each arm.
  6. Outcomes in each arm.
  7. Compare the two sets of outcomes.

This same diagram is reused later to anchor where in the trial process bias is minimised.

Context, regulation and guideline documents

  • Clinical trials are experiments in patients designed to evaluate the benefits and risks of treatment(s).
  • Regulatory authorities require rigorous scientific and ethical processes for studies supporting applications for licensing of drugs and devices. The bodies shown are the FDA, MHRA, EMEA and Medsafe (NZ Medicines and Medical Devices Safety Authority).
  • Clinical guideline groups, the Cochrane Collaboration, and discipline specific bodies such as the National Cancer Institute also play their part.
  • Guideline documents:
    • ICH, the International conference on harmonisation, from drug regulatory authorities and the pharmaceutical industry.
    • CONSORT, guidelines for reporting of randomized controlled trials (MRC, Cancer Research UK and others), “Transparent Reporting of Trials”.
    • Equator network, funded by public good such as NHS and MRC, “Enhancing the QUAlity and Transparency Of health Research”.
    • SPIRIT, guidelines for writing trial protocols, Canadian public good funding, “Standard Protocol Items: Recommendations for Interventional Trials”.

Ethics of phase III trials

  • Reliable answer to a relevant question: the research question should be relevant to the population in which the trial is carried out, the design should be sound, and the trial should be carried out according to good practice guidelines (ICH, CONSORT).
  • Individual ethics: equipoise and informed consent, safety of participants (minimise harm and maximise benefit), and independent oversight by a Data Monitoring Committee.
  • Group ethics: approve new beneficial treatments as rapidly as possible, and avoid approving ineffective or harmful treatments.

Research questions: superiority and non-inferiority

Two common scenarios:

  • Superiority: evaluate whether a new treatment is superior to an existing treatment or to no treatment. Example: do statins reduce mortality in people with cardiovascular disease?
  • Non-inferiority: evaluate whether a new treatment is no worse than an existing treatment. Example: is recurrence free survival as good with laparoscopic surgery as with open surgery for colon cancer?

Superiority example: the LIPID study

LIPID Study Group, “Prevention of cardiovascular events and death with pravastatin in patients with coronary heart disease and a broad range of initial cholesterol levels”, NEJM 1998;339(19):1349-57.

  • Patients with a history of myocardial infarction or hospitalization for unstable angina, and initial plasma total cholesterol 155 to 271 mg per decilitre.
  • 9014 patients randomised to pravastatin 40 mg daily or placebo.
  • Mean follow-up 6.1 years.
  • Time to event outcome, event = CHD death. The sample size was based on the number of events.
  • Result: RR = 0.76, 95% CI (0.65 to 0.87), which translates to a 24% risk reduction, 95% CI (13% to 35%), P<0.001. The Kaplan–Meier curves show cumulative risk of death due to CHD reaching about 10% on placebo versus about 7% on pravastatin. (The later slide restating this study gives the confidence interval for the risk reduction as 12% to 35%.)

Non-inferiority example: the COST study

Clinical Outcomes of Surgical Therapy Study Group, “A Comparison of Laparoscopically Assisted and Open Colectomy for Colon Cancer”, NEJM 2004;350:2050-9.

  • Patients with adenocarcinoma of the colon due to undergo colectomy.
  • 872 patients randomised to open or laparoscopically assisted colectomy, performed by credentialed surgeons.
  • Median follow-up 4.4 years. Primary end point was recurrence free survival.
  • RR of recurrence free survival 0.86, 95% CI (0.63 to 1.17). RR of overall survival 0.91, 95% CI (0.68 to 1.21). Cumulative incidence of recurrence curves lie close together, both reaching roughly 0.2 by year 5, P=0.32.
  • Short-term advantages of laparoscopic surgery: median hospital stay five days versus six days (P<0.001), use of parenteral narcotics three days versus four days (P<0.001), oral analgesics one day versus two days (P=0.02).

Measurement of treatment effects

  • Two things must be decided: the outcome measures and the time frame.
  • For a phase III trial the outcome measures needed are a measure of efficacy that is tangible to patients (length of survival, quality of life) and measures of safety.
  • Outcome measures must be able to be measured on everyone, and must be validated and reliable.

Biomarkers

  • Biomarkers are biological measures of the disease process, or of the effect of treatment on the disease process.
  • Traditional biomarkers by field: oncology, tumour response. HIV/AIDS, CD4 count and viral load. Cardiovascular disease, lipids and arrhythmia. Vaccines, immune response.
  • Studies with a biomarker as an endpoint can establish that there is a biological effect, but they do not necessarily tell us about clinical efficacy.
  • Three examples where the biomarker improved but the clinical outcome did not follow:
    • Clofibrate: successful in improving lipid levels but no impact on mortality.
    • AZT: effective in improving immune function in people with HIV infection, but did not prolong survival.
    • Anti-arrhythmics (encainide, flecainide): successful in suppressing arrhythmias but increased risk of sudden death.

Surrogate endpoints

  • A surrogate endpoint is a variable which can be used in place of the clinical outcome to provide definitive evidence of treatment efficacy in a phase III trial, and it must fully capture the effect of treatment on the clinical outcome.
  • The setting giving the greatest potential for a surrogate to be valid, along a time axis: the intervention acts on the disease, the disease leads to the surrogate end point, and the surrogate end point leads to the true clinical outcome. The surrogate is only valid when it lies fully on the causal pathway between disease and true clinical outcome.
  • Validity is hard to demonstrate, needs to be done for the specific treatment and disease setting, and requires several large trials.
  • Biomarkers which have not been demonstrated to be surrogates can be useful primary outcome measures at phase II but not at phase III.

Duration of follow-up

  • Laparoscopic versus open surgery for colon cancer: non-inferiority on survival, early quality of life benefits anticipated, survival differences may not be apparent until 5 to 7 years after treatment, and follow-up must be long enough to rule out safety risks.
  • Bone marrow transplantation: early harm (graft versus host disease) with long term survival benefit, so follow-up must be long enough to see whether benefits outweigh risks.
  • Contrasting case of a rapid, visible effect: vemurafenib in BRAF mutant melanoma refractory to standard therapy produced a dramatic reduction in the size of multiple subcutaneous nodules after 15 weeks of treatment (Wagle N, J Clin Oncol 2011;29:3085-3096).

Minimising bias: the three-part framework

Minimising bias in the design:

  • Method of allocation (randomisation).
  • Concealment of allocation (blinding).
  • Assessment methods that minimise bias (blinding).

Minimising bias during the trial:

  • Adherence to intervention and control.
  • Adherence to protocol.
  • High levels of retention and follow-up.
  • Complete and accurate capture of outcome data.

Minimising bias in the analysis and reporting:

  • Intention to treat analysis.
  • Limited number of statistical analyses.
  • Confirmatory versus exploratory analyses.
  • Publication of the trial protocol.
  • Pre-specification of appropriate statistical analyses.

Randomisation

  • Aim: to divide patients into two (or more) groups with the same distribution of characteristics, which controls confounding, so that at the end of the trial any differences in outcome between the groups can be attributed to differences between the treatment and the control.
  • Randomisation is the best tool we have for controlling confounding, but simple randomisation can still leave imbalance between groups, particularly in smaller trials.
  • Stratification and blocking are used to force balance on chosen factors, for example clinical site, a prognostic factor, or time.
  • The balance achieved by randomisation must be maintained through to measurement of outcome and analysis.

Concealment of allocation

  • Concealment of allocation: the person recruiting patients to the trial should not be able to guess which arm the patient will be in. Schulz and Grimes (Lancet 2002;359:614-18) portray attempts to defeat it as “deciphering the allocation concealment scheme”.
  • Worked example, American Heart Association trial of anticoagulant therapy: control group were patients admitted to hospital on even days, treatment group those admitted on odd days. Patients on anticoagulant fared much better, but there were far more patients on anticoagulant than on control, which is the clue that allocation could be anticipated.
  • The same problem in a published trial: Bernard et al, “Treatment of Comatose Survivors of Out-of-Hospital Cardiac Arrest with Induced Hypothermia”, NEJM 2002;346:557-563. 77 patients in Melbourne between September 1996 and June 1999 were “randomly assigned” to hypothermia (core body temperature reduced to 33°C within hours of the return of spontaneous circulation and maintained for 12 hours) or normothermia, but assignment was according to the day of the month, with hypothermia on odd-numbered days. Primary outcome was survival to hospital discharge with sufficiently good neurologic function to be discharged home or to a rehabilitation facility. Inclusion required an initial rhythm of ventricular fibrillation on ambulance arrival, successful return of spontaneous circulation, persistent coma afterwards, and transfer to one of four participating emergency departments. Exclusions were age under 18 for men, age under 50 for women (possibility of pregnancy), cardiogenic shock (systolic BP under 90 mmHg despite epinephrine infusion), possible causes of coma other than cardiac arrest (drug overdose, head trauma, cerebrovascular accident), and no available intensive care bed. Day-of-month allocation is quasi-randomisation, so allocation could be anticipated or deduced ahead of time.
  • To conceal allocation:
    • Use placebos and blinding whenever possible.
    • Do not let anyone involved in recruiting patients see the randomisation lists. Web-based randomisation systems are good.
    • Do not tell people the details of the stratification and blocking used. Random block sizes are helpful.

Placebos

  • Placebos are biologically inactive substances or procedures used to make the treatment and control regimens look the same, for example look-alike pills or sham surgery.
  • They cannot always be used because of treatment side effects, ethics (risk), science (HIV vaccine trials), or the availability of an inert placebo (for example microbicide trials).
  • Placebos can be used with both no-treatment comparators and active control comparators.

Blinding

  • Blinding (or masking) means the person who is blinded is not aware of whether patient allocation is to treatment or control.
  • The single/double blind terminology is confusing, so it is better to think about who is blinded: the patient, the treating physician, nursing care, or the person measuring the outcome.
  • Ideal: everyone is blinded during conduct of the trial, unless unblinding with respect to a particular patient is required for safety reasons.

Adherence

  • To allocated treatment: patients have the right to withdraw from their allocated treatment at any time. Adherence to the allocated treatment and control is important for obtaining a good estimate of treatment efficacy, and poor adherence, for example due to toxicity or patient choice, will limit the effectiveness of treatment.
  • To protocol: deviations from protocol in the timing or methods of assessment often lead to bias, and also affect trial credibility.

Retention and completeness of follow-up

  • Patients have the right to withdraw from a trial at any time.
  • Aim for 95% complete follow-up.
  • It is important to encourage patients who withdraw from treatment to remain in follow-up.
  • Small amounts of data missing at random can be dealt with in the analysis, but missing data often reflect the health of the patient, for example being too sick to fill out a form. These are non-ignorable missing data, nothing can be done about them after the fact, and they will cause bias.
  • There are particular problems with bias if follow-up differs between the treatment and control groups.

Intention to treat analysis

  • The problem: the study population splits into experimental treatment and control, but each arm generates non-compliers, and it must be decided how those patients are analysed.
  • Forms of non-adherence to protocol-allocated treatment:
    • Discontinue experimental treatment.
    • Discontinue control treatment.
    • Contamination or cross over: experimental treatment to control, or control to experimental treatment.
  • Examples: the control group accessing the experimental treatment outside the trial (for example nutritional supplements such as vitamin D, or lifestyle interventions), an early HIV trial, and laparoscopic surgery crossing over to open surgery. In the COST study, 25% of those randomised to laparoscopic surgery received open surgery. Oncology trials are another setting, where patients often withdraw from treatment due to adverse events or disease progression.
  • Under intention to treat analysis, patients randomised to laparoscopic surgery who are converted to open surgery are still counted in the laparoscopic-assigned outcome group.
  • Intention to treat analysis is the primary analysis for superiority trials. It compares outcome in all those allocated to experimental treatment with all those allocated to control, and:
    • requires no exclusions post randomisation, complete follow-up, and complete assessment of outcome;
    • preserves the balance between treatment and control groups achieved by randomisation, that is, control of confounding;
    • allows for real world non-adherence (acceptability of treatment, adverse effects);
    • provides an evaluation of the effectiveness of treatment.
  • Efficacy asks whether the treatment works in an idealised setting, and can only be estimated without bias in a trial with near complete adherence. Effectiveness investigates performance in a real-world setting, where not all patients are adherent.

Important

Coronary Drug Project mortality rates show why analysing by adherence rather than by allocation misleads. Clofibrate: 18.2% overall, 15.0% in those with >80% adherence, 24.6% in those with <80% adherence. Placebo: 19.4% overall, 15.1% with >80% adherence, 28.2% with <80% adherence. Adherent patients did better in both arms, including on placebo, so the apparent benefit of adherence is not a treatment effect.

Multiplicity of statistical analyses

  • Statistical analysis provides an estimate of the effect of treatment, a confidence interval (a range of plausible values for the true effect of treatment in the target population), and a p-value (the probability we would have seen an effect this large or larger if the treatment did not in fact work).
  • Worked figures from the LIPID study of pravastatin in the prevention of mortality in patients with heart disease: estimated risk reduction 24%, 95% confidence interval (12% to 35%), P-value <0.001.
  • In many trials people carry out multiple statistical analyses, each with its own confidence interval and p-value. The problems are that the effects of random variation on the results are bigger than people generally think, and that the interpretation of confidence intervals and p-values depends on how many analyses have been done.
  • The p-value in detail: we control the chance of getting a single false positive result from a trial by setting the significance level, usually 0.05. If a p-value is <0.05 we consider we have evidence of a treatment effect in this hypothesis testing framework, and a 95% confidence interval will exclude the null value. With many tests, the chance of a false positive across all the tests is the chance that any one or more gives a p-value <0.05, which is much higher than 5%: with 20 tests it is over 50%. This is what “data dredging”, “fishing” and “p-hacking” refer to, since if you look hard enough you will find something.
  • Two problems with multiple statistical tests:
    • Inflation of the chance of false positive results, which analyses need to adjust for.
    • Potential over-estimate of treatment benefit due to selective reporting of the analyses which appear favourable to treatment, that is, “random highs”.
  • Quoted caution: “It ain’t so much the things we don’t know that get us into trouble. It’s the things we know that ain’t so.” (Artemus Ward, 1834-1867)

Repeated analyses during a trial

  • Repeated looks at the data during a trial are desirable if done appropriately, because they allow early termination of the trial if the questions are answered reliably earlier than anticipated. But care is needed.
  • MRC AML12 compared 4 versus 5 courses of therapy in patients with acute myeloid leukaemia, asking whether an extra course of consolidation therapy confers additional benefit. The second interim analysis gave results in favour of 5 versus 4 courses (Wheatley et al, Haematology 2002):
    • 1997: five courses 7/102 deaths, four courses 15/100, O-E −4.6, variance 5.5, odds reduction 57% (SD 29), 2P = 0.05.
    • 1998 (1): five courses 23/171, four courses 42/169, O-E −12.0, variance 15.9, odds reduction 53% (SD 18), 2P = 0.003.
    • Both point estimates fell on the “five courses better” side of 1.0.

Subgroup analyses

  • ISIS-2, “Randomised trial of intravenous streptokinase, oral aspirin, both, or neither among 17187 cases of suspected acute myocardial infarction” (Lancet 1988). Overall result: 804/8587 (9.4%) vascular deaths on aspirin versus 1016/8600 (11.8%) on placebo tablets, a 23% (SD 4) odds reduction.
    • Subset analysis by prior MI: prior MI yes, RR = 1.02, so aspirin showed almost no benefit; prior MI no, RR = 0.74, an apparent benefit. Same underlying trial.
    • Subset analysis by astrological birth sign: Gemini and Libra RR = 1.1, other birth signs RR = 0.74. This deliberately absurd example shows that subgroup analyses can produce spurious “significant” differences purely by chance when many subgroups are tested.
  • Actimmune (interferon gamma-1b) in idiopathic pulmonary fibrosis, a life-threatening disease with no known effective therapy (Fleming TR, Ann Intern Med 2010;153(6):400-406):
    1. A randomized placebo-controlled trial of 330 patients was conducted, sponsored by InterMune, with progression-free survival as the primary end point.
    2. Data shown to the independent Data Monitoring Committee in June 2002, the prespecified trial end date: Actimmune n=162 with 75 events (46.3%) versus placebo n=168 with 87 events (51.8%), p=0.53, that is, not statistically significant.
    3. 28 August 2002 sponsor news release: “InterMune Announces Phase III Data Demonstrating Survival Benefit of Actimmune in IPF…Reduces mortality by 70% in patients with mild to moderate disease, (p = 0.004)…. The mortality benefit is very compelling and represents a major breakthrough in this difficult disease.” This came from a post-hoc subgroup, not the prespecified primary analysis.
    4. Final published results (Raghu G et al, NEJM 2004;350:125-33) made no mention of the subgroup.
    5. The confirmatory trial (King TE Jr et al, INSPIRE, Lancet 2009;374:222-28) was stopped early with no evidence of benefit.
    6. The former InterMune CEO was sentenced for false and misleading statements related to the drug’s clinical tests (US Department of Justice, 14 April 2011), and appealed the fraud conviction over interpretation of results (Nature 2013;502:17-18).

Warning

The transcript flags slide 58 (the sponsor news release slide) as not rendered, so the layout and emphasis of the ”?” on that slide is not verifiable from text extraction.

References given for the bias and reporting material

  • Schulz KF, Chalmers I, Hayes RJ, Altman DG. Empirical evidence of bias: dimensions of methodological quality associated with estimates of treatment effects in controlled trials. JAMA 1995;273:408-12.
  • Schulz KF, Grimes DA. Allocation concealment in randomised trials: defending against deciphering. Lancet 2002;359:614-18.
  • Hollis S, Campbell F. What is meant by intention-to-treat analysis? Survey of published randomised controlled trials. BMJ 1999;319:670-4.
  • Chan AW, Hrobjartsson A, Haahr MT, Gotzsche PC, Altman DG. Empirical evidence for selective reporting of outcomes in randomized trials: comparison of protocols to published articles. JAMA 2004;291:2457-65.
  • Fleming TR. Clinical Trials: Discerning Hype From Substance. Ann Intern Med 2010;153(6):400-406.

Critical appraisal: what and why

  • Critical appraisal is part of the third step of evidence-based practice, appraising the evidence for validity and clinical usefulness.
  • “Critical” is meant in the sense of a systematic and thorough review of a paper: there may be a lot to criticise, or a lot to praise.
  • To do it you need to know how such a study should have been done, and where and how things might go wrong in such a study.
  • Doctors need these skills because of the constant evolution of evidence about causes of diseases, best treatments and preventive measures.
  • You cannot unthinkingly accept the results of a study just because it was published in a high impact journal, had statistically significant findings, was a huge study, came from a prestigious hospital or university, or was a randomised controlled trial. Appraisal skills need to be developed and practised.
  • Relevance for this course: three EBP tutorials, the end of year exam, and the rest of your career.

Reporting guidelines by study type (EQUATOR library)

  • Randomised trials: CONSORT.
  • Observational studies: STROBE.
  • Systematic reviews: PRISMA.
  • Study protocols: SPIRIT.
  • Also listed: case reports (CARE), clinical practice guidelines (AGREE/RIGHT), qualitative research (SRQR/COREQ), animal pre-clinical studies (ARRIVE), quality improvement studies (SQUIRE), economic evaluations (CHEERS).

General anatomy of a paper

Top to bottom: title, abstract, main body of the paper, acknowledgements, references.

  • The main body follows IMRAD: Introduction, Methods, Results, And Discussion.
  • The abstract is useful for screening.
  • Acknowledgements cover the funding source and conflicts of interest.

Reminder: the GATE picture

GATE is the Graphic Approach to Epidemiology/EBP. Its elements:

  • Population (problem, patients, participants), drawn as the inverted triangle “P”.
  • Exposure (intervention), the “E(I)” part of the circle.
  • Comparison (control), the “C” part of the circle.
  • Outcomes, the square “O”.
  • Time, the “T” arrow, indicating that outcomes are assessed over time following exposure/intervention versus comparison in a population.

Introduction

Should include:

  • The rationale for undertaking the study (importance of the exposure and/or outcome of interest).
  • What was, and was not, already known.
  • The a priori hypothesis or study question.
  • A clear statement outlining both the type and the aims of the study.

Methods

Should include:

  • The type of study.
  • The setting and relevant dates.
  • The participants.
  • The key exposure and outcome, plus other important variables such as potential confounders, and how they were measured.
  • How the study size was decided upon.
  • Statistical methods (analysis).

Think about:

  • Whether the study design was appropriate.
  • Whether the study had sufficient power.
  • Methods to minimise bias.
  • Methods to minimise confounding.
  • Is it worthwhile reading any further?

Populations: source, eligible, participating

Three nested levels, narrowing successively:

  1. Source population/setting: the population from which the participants were drawn.
  2. Eligible population: people from the source population who met the eligibility criteria and did not meet any exclusion criteria.
  3. Participating (sample) population: people who met the eligibility criteria and agreed to take part in the study.

Sometimes the eligible population equals the participating population, for example studies based on existing health records when all records are found.

Results

Should include:

  • Numbers of individuals at each stage of the study, reported as a CONSORT flow diagram.
  • Characteristics of participants (Table 1).
  • Key findings related to the a priori question(s) and a priori analysis plan.
  • Results of any sensitivity analyses.

Worked examples of the first two requirements: the ALLHAT-LLT CONSORT diagram (JAMA Intern Med 2017;177(7):955-965) and the baseline characteristics table from the rosuvastatin versus placebo trial (NEJM 2016;374:2021-31), which compares the two arms on age, sex, cardiovascular risk factors (waist-hip ratio, smoking, HDL cholesterol, glucose tolerance, diabetes, family history, renal dysfunction, hypertension), number of risk factors, blood pressure, heart rate, BMI, cholesterol (total, LDL, HDL), triglycerides and fasting glucose.

Reminder: the square (occurrence and effect estimates)

Based on the 2x2 exposed (yes/no) by outcome (yes/no) contingency table with cells a, b, c, d.

  • Categorical outcomes: occurrence estimates are the incidence in the exposed group and in the comparison group; effect estimates are the relative risk and the absolute risk difference.
  • Numerical (continuous) outcomes: occurrence estimates are the mean, or the mean change from baseline, in the exposed group and in the comparison group; effect estimates are the difference in means or in mean changes.

Discussion

Should include:

  • Statement of principal findings.
  • Strengths and limitations of the research.
  • Results in context, that is, findings in relation to other studies.
  • Meaning, for example possible mechanisms, generalisability, implications for clinicians and/or policymakers.
  • Unanswered questions and future research.

Finally

It is important to ascertain the funding source(s) and the role of the funder (for example any role in the design, conduct or reporting), and whether the authors had any other conflicts of interest.

Back to Mr X: the evidence on statins in older adults

Framing the decision

Clinical decisions are a balance, with patient preference at the fulcrum:

  • Benefit pan: primary prevention of cardiovascular events in people aged over 75 years.
  • Risk pan: risk of myopathy and other adverse reactions in people aged over 75 years, with risk increasing in the presence of various drugs and co-morbidities; and the risks of polypharmacy and overprescribing.

The literature is actively contested, shown as paired pro/con reviews: Razavi, Mehta and Sperling, “Statin therapy for the primary prevention of cardiovascular disease: Pros”, Atherosclerosis 356 (2022) 41-45; Durai and Redberg, “Statin therapy for the primary prevention of cardiovascular disease: Cons”, Atherosclerosis 356 (2022) 46-49; and an editorial, “Awaiting the verdict: Statins and the road ahead for primary prevention in older adults”, Journal of the American Geriatrics Society.

Two RCTs used to frame the PICOT question

ALLHAT-LLT (Han et al, “Effect of Statin Treatment vs Usual Care on Primary Cardiovascular Prevention Among Older Adults”, JAMA Intern Med 2017;177(7):955-965):

  • CONSORT flow: 10355 ALLHAT-LLT cohort (5170 pravastatin, 5185 usual care) → 4546 excluded for age under 65 years → 5809 aged 65 years and over (2913 pravastatin, 2896 usual care) → 2942 excluded for evidence of ASCVD at baseline → 2867 with no ASCVD at baseline (1467 pravastatin, 1400 usual care).
  • Primary outcome all-cause mortality, pravastatin versus usual care: overall HR 1.18 (0.97 to 1.42), p=0.09; age 65-74 years HR 1.08 (0.85 to 1.37), p=0.55; age 75 years and over HR 1.34 (0.98 to 1.84), p=0.07. Pravastatin showed numerically higher mortality than usual care over 6 years, more pronounced in the 75 and over subgroup, but none reached conventional statistical significance.
  • Stroke (fatal and nonfatal), pravastatin versus usual care: overall HR 1.06 (0.76 to 1.49), p=0.72; 65-74 years 1.03 (0.68 to 1.57), p=0.89; 75 years and over 1.09 (0.63 to 1.90), p=0.76.

HOPE-3 (Yusuf et al, “Cholesterol Lowering in Intermediate-Risk Persons without Cardiovascular Disease”, NEJM 2016;374:2021-31):

  • CONSORT flow for the rosuvastatin versus placebo comparison: 14,682 included in run-in → 1,977 (13.5%) excluded (509 (3.5%) side effects, 844 (5.7%) compliance under 80%, 483 (3.3%) unwilling to continue, 141 (1.0%) other reasons) → 12,705 randomized (86.5%), 6,361 assigned to rosuvastatin and 6,344 to placebo → primary outcome status ascertained in 6,308 and 6,284 respectively (45 lost to follow-up in each arm, 8 withdrew consent on rosuvastatin, 15 on placebo) → 6,361 and 6,344 included in the analysis, with 0 excluded from analysis in either arm.
  • First coprimary outcome, a composite including death from cardiovascular causes, occurred in 235 participants (3.7%) in the rosuvastatin group and 304 (4.8%) in the placebo group: hazard ratio 0.76, 95% CI (0.64 to 0.91), P=0.002.
  • Other outcomes, all favouring rosuvastatin: second coprimary outcome HR 0.75 (95% CI 0.64 to 0.88), P<0.001; stroke HR 0.70 (0.52 to 0.95), P=0.02; myocardial infarction HR 0.65 (0.44 to 0.94), P=0.02; coronary revascularization HR 0.63 (0.44 to 0.91), P=0.01.

RCTs currently underway

  • STAREE, “STAtins in Reducing Events in the Elderly” (Monash University). Protocol: “Statins for extension of disability-free survival and primary prevention of cardiovascular events among older people: protocol for a randomised controlled trial”, BMJ Open 2023;13(4):e069915.
  • PREVENTABLE, “Pragmatic evaluation of events and benefits of lipid lowering in older adults: trial design and rationale”, Journal of the American Geriatrics Society. Its promotional material states that about 1 in 3, or 30%, of people without heart disease are still taking a statin after their 75th birthday, and that PREVENTABLE will help us understand if that makes sense.

What to do

Evidence-based practice is the central overlap of three circles: patient values and choices, best available research evidence, and clinical expertise. This is the framework offered for deciding about Mr X’s statin. Systematic reviews and meta-analyses are covered in a separate lecture on 27 May.

Key points

  • An RCT is the best study design to test whether a treatment or preventive intervention is beneficial.
  • But you cannot rely on the results of a study just because it is an RCT, because they can be done badly.
  • Doctors need to know how RCTs should be done, and know how to critically appraise papers that report on RCTs.

Self-test

  1. List what phases I and II of treatment evaluation screen for, state what phase III trials provide, and name the design phase III trials always use.
  2. Describe the stages of a randomised controlled trial in order, from patient population through to the comparison of outcomes.
  3. List the three requirements the lecture sets for an evaluation of an intervention, and the individual and group ethical requirements for a phase III trial.
  4. Distinguish a superiority trial from a non-inferiority trial, giving the lecture’s clinical example of each.
  5. The LIPID study reported RR 0.76, 95% CI (0.65 to 0.87) for CHD death. Translate this into a risk reduction with its confidence interval, and state what the sample size was based on.
  6. What did the COST trial find for recurrence free survival and overall survival, and what short-term outcomes favoured laparoscopic colectomy?
  7. What properties must a phase III outcome measure have, and what two kinds of outcome measure are needed?
  8. Define a biomarker, give the traditional biomarker used in each of oncology, HIV/AIDS, cardiovascular disease and vaccines, and explain the limitation of a biomarker endpoint.
  9. Give the three examples of treatments that improved a biomarker without delivering clinical benefit, and state what happened to the clinical outcome in each.
  10. Define a surrogate endpoint, state the structural condition under which a surrogate can be valid, and say at which trial phase an unvalidated biomarker may still serve as a primary outcome.
  11. Explain why the duration of follow-up must be chosen carefully, using the colon cancer surgery and bone marrow transplantation examples.
  12. Describe the three-part framework for minimising bias in an RCT, listing the components of each part.
  13. Explain the aim of randomisation, what it controls, and why stratification and blocking are added to it.
  14. Define concealment of allocation, and explain using the anticoagulant and induced hypothermia examples how allocation by day of admission breaks it.
  15. List the ways of concealing allocation.
  16. Define a placebo and list the reasons a placebo cannot always be used.
  17. Define blinding, list who might be blinded, and state the ideal.
  18. Distinguish adherence to allocated treatment from adherence to protocol, and say why each matters.
  19. What level of follow-up completeness should a trial aim for, and what are non-ignorable missing data?
  20. List the forms non-adherence to protocol-allocated treatment can take, with an example of contamination.
  21. Define intention to treat analysis, and state what it requires, what it preserves, and what it evaluates.
  22. Distinguish efficacy from effectiveness, and say which can only be estimated without bias in a trial with near complete adherence.
  23. In the COST study, what proportion of patients randomised to laparoscopic surgery received open surgery, and how are those patients handled under intention to treat?
  24. Using the Coronary Drug Project mortality figures, explain why analysing by adherence rather than by allocation is misleading.
  25. What three things does a statistical analysis provide, and what does each one mean?
  26. Explain why doing many statistical tests inflates the chance of a false positive, give the figure for 20 tests, and name the informal terms used for this practice.
  27. Describe the two problems the lecture identifies with multiple statistical tests.
  28. Why are repeated looks at the data during a trial desirable, and what does the MRC AML12 example show about them?
  29. Using ISIS-2, explain why subgroup analyses should be treated with caution.
  30. Describe the sequence of events in the Actimmune/idiopathic pulmonary fibrosis case and state what it illustrates.
  31. What is critical appraisal, which step of evidence-based practice does it belong to, and what two things do you need to know in order to do it?
  32. List the reasons the lecture gives for why a study’s results cannot be accepted uncritically.
  33. Name the reporting guideline for each of randomised trials, observational studies, systematic reviews and study protocols.
  34. Describe the general anatomy of a paper, including what IMRAD stands for, and say which part is useful for screening.
  35. List the elements of the GATE picture and what the time element indicates.
  36. What should the Introduction of a paper include?
  37. List what the Methods section should include, and the questions to ask of it when appraising.
  38. Distinguish the source, eligible and participating populations, and give a case in which the eligible population equals the participating population.
  39. What should the Results and Discussion sections include?
  40. For categorical and for numerical outcomes, name the occurrence estimates and the effect estimates.
  41. What should you ascertain about funding and conflicts of interest, and why?
  42. A trial reports a non-significant primary outcome but a striking benefit in one subgroup. Predict how that finding is likely to behave in a confirmatory trial, and explain your reasoning.
  43. Using the Ministry of Health CVD risk table, state the drug therapy recommendation and the follow-up interval for each of the four risk bands.
  44. What benefit and what risks sit on the two pans of the balance for a statin decision in a man aged over 75, and what sits at the fulcrum?
  45. Summarise the ALLHAT-LLT results for all-cause mortality and for stroke, and state the direction of effect.
  46. Summarise the HOPE-3 first coprimary outcome result, and describe what the run-in period did to the randomised population.
  47. Name the two ongoing or recent trials of statins in older adults, and state what PREVENTABLE’s promotional material says about current prescribing.
  48. Name the three components of evidence-based practice and how they relate, then apply them to Mr X: explain how you would approach the decision whether to restart his statin, drawing on the trial evidence, the appraisal points and his stated values.

Answers