Critical Appraisal and Evidence-Based Medicine for Emergency Clinicians

EDUCATIONAL RESOURCEThis page teaches appraisal method. It is not clinical guidance and must not be used to make a treatment decision — the trials quoted here are worked examples, not recommendations.
ResusDoc
Critical Appraisal
Evidence-based medicine · method, statistics & worked papers
Main site

Appraisal is a clinical skill, not a statistics exam

The question a clinician actually needs to answer about a paper is narrow: would acting on this change what I do to the next patient through my door, and how badly could it be wrong? Almost everything below exists to answer that. The statistics are a means to it, not the point of it.

Three habits do most of the work. Read the methods before the results — the design fixes the ceiling on what the numbers can mean, and by the time you reach the abstract's conclusion you are being told what to think. Convert every relative number into an absolute one, because a 30 % risk reduction is a different intervention at a 20 % baseline risk than at a 0.2 % one. And ask what the trial's population was, because population mismatch, not statistics, is the commonest reason a true result does not apply to your patient.

The twenty-three chapters here run from describing a single variable to GRADE. The Calculators tab does the arithmetic, the Figures tab shows what the plots in a results section are hiding, the Worked papers tab appraises six real trials, each chosen because it teaches one trap cleanly.

A paper read without a question in mind is a paper that will tell you its own conclusion. Writing the question down first is not a formality — it is what makes the relevance judgement happen before you have been persuaded.

ElementAsksWhere appraisal goes wrong
PopulationWhich patients?The trial's population is narrower than the specialty that adopts it. Severity, age, setting and time-from-onset are the usual mismatches.
InterventionExactly what, at what dose, how soon?A regimen gets simplified in the retelling. A 1 g bolus plus an 8-hour infusion becomes "give tranexamic acid".
ComparatorAgainst what?Placebo answers "does it work"; usual care answers "should I switch". A trial against a straw-man comparator flatters the drug.
OutcomeWhich outcome, measured when?The outcome that made the headline is often not the primary outcome, and often not one a patient would notice.
Background versus foreground. Background questions ("how does this drug work") are answered by textbooks. Foreground questions ("in adults with GI bleeding, does tranexamic acid reduce death from bleeding compared with placebo") are answered by studies, and only foreground questions can be appraised. If you cannot phrase your question in PICO, you are not yet ready to judge the paper — you are still reading to learn the topic, which is a different and perfectly respectable activity.

Then predict the answer. Before you read the results, commit to what you expect and how big an effect would change your practice. This single move protects against the two commonest failures at once: being impressed by a result you would have dismissed had it gone the other way, and calling a trial "negative" when it merely failed to find the implausibly large effect it was powered for.

The tree

The first fork is interventional versus observational: did the investigators allocate the exposure, or only watch it? Everything about what a study can claim follows from that. The second fork, within observational work, is the direction of sampling: did you sample on exposure and wait for outcomes (cohort), or sample on outcome and look back for exposure (case–control)?

DesignSamplingAnswers wellStructural weakness
Randomised controlled trialAllocated by chanceDoes this intervention cause this outcomeCost, narrow eligibility, often short follow-up, rarely powered for rare harms
Cluster RCTGroups randomisedSystems, pathways, trainingEffective sample size sits between the number of clusters and the number of patients, and is nearer the former the larger the clusters and the more alike patients within them; needs ICC adjustment in both the sample size and the analysis
Stepped-wedgeSequential rolloutRoll-outs that cannot be withheldConfounded by secular trend; needs time in the model
Prospective cohortOn exposureIncidence, prognosis, multiple outcomesConfounding; loss to follow-up
Retrospective cohortOn exposure, from recordsRare exposures, long latencyData collected for another purpose; exposure misclassification
Case–controlOn outcomeRare outcomes, many exposuresRecall and selection bias; gives odds ratios, not risks
Cross-sectionalOne time pointPrevalence, diagnostic accuracyNo temporality — cannot separate cause from consequence
Case series / reportWhoever appearedNovelty, harm signals, the unexpectedNo denominator, no comparison — cannot estimate effect at all
Diagnostic accuracyConsecutive suspected casesHow well a test classifiesReference-standard quality; spectrum and selection effects
Systematic reviewStudies, not patientsSynthesis, precision, consistencyInherits its constituents' bias; heterogeneity may be real
The hierarchy is a heuristic, not a law

The pyramid with RCTs near the top is a statement about average susceptibility to bias, not a ranking of individual papers. A small, unblinded, selectively reported RCT stopped early for benefit is weaker evidence than a large, carefully adjusted prospective cohort with complete follow-up. Judge the study, not its label.

Where the hierarchy inverts in practice. Some questions cannot be randomised — you cannot allocate a patient to strangulation, to a toxic ingestion, or to a blinding eye injury. For those, the best available evidence is observational by necessity, and demanding an RCT is not rigour but paralysis. Say so explicitly rather than pretending the evidence is stronger than it is; the honest sentence is "the recommendation is strong and the evidence is thin, and it will stay thin."
Pragmatic versus explanatory

An explanatory trial asks whether a thing can work under ideal conditions: tight eligibility, protocol-driven care, expert centres. A pragmatic trial asks whether it does work in the real world: broad eligibility, usual care as comparator, routinely collected outcomes. Pragmatic trials generalise better and tend to show smaller effects — which is usually the truth rather than a failure. When a pragmatic trial disappoints after an explanatory one impressed, the effect has not disappeared; it has been measured under conditions you actually work in.

Before any test is chosen or any p-value computed, a paper has to say what its data are. Almost every later decision — which summary to report, which test is legitimate, whether a mean is even meaningful — follows from the answer.

What kind of data is it?
TypeExampleLegitimate summaryWatch for
Nominal (categorical, unordered)Blood group, presenting complaint, discharge destinationCounts and proportionsThere is no average. A "mean blood group" is not a quantity.
BinaryDied / survived, admitted / dischargedProportion, risk, oddsA special case of nominal, and the one most of this site is about.
Ordinal (ordered, unequal gaps)Triage category, pain score 0–10, mRS, GCSMedian and IQR; proportions above a thresholdThe gaps are not equal — the step from mRS 4 to 5 is not the step from 0 to 1. Means of ordinal scales are reported constantly and are hard to interpret.
Discrete countAttendances per shift, number of seizuresMedian and IQR, or a rateOften skewed and bounded at zero.
ContinuousAge, lactate, door-to-CT time, sodiumMean and SD if symmetric; median and IQR if notMeasurement precision, and whether it was later chopped into categories.
Dichotomising a continuous variable throws information away. Turning lactate into "high vs normal" or age into "over 65 vs under" is convenient and costs statistical power, hides a dose–response relationship, and makes the result depend on a cut-off someone chose. When a paper reports a continuous predictor only as a binary split, ask where the cut-off came from — and whether it was chosen after looking at the data, which makes the reported performance optimistic.
Measures of the middle
  • Mean — the arithmetic average. Uses every value, which is its strength and its weakness: one extreme value drags it. Meaningful only for interval or ratio data.
  • Median — the middle value when ordered. Unaffected by how extreme the extremes are, so it survives skew and outliers. The right choice for most emergency-department time data.
  • Mode — the commonest value. Rarely useful for continuous data, but it is the only measure of the middle available for nominal data, and it is what matters when a distribution has two peaks.

In a perfectly symmetric distribution the three coincide. The gap between mean and median is itself a measure of skew — if a paper reports both and they differ substantially, the data are skewed whatever else the text says.

Measures of spread
MeasureWhat it isUse when
RangeLowest to highestAlmost never on its own — it depends entirely on the two most extreme patients and grows with sample size
Interquartile range (IQR)25th to 75th centile — the middle half of the dataSkewed data, ordinal data, anything with outliers. Report it as two numbers (Q1–Q3), not as their difference alone
VarianceMean squared deviation from the meanRarely reported directly — its units are squared. The worked dataset on the Figures tab has an SD of 44.7 minutes and therefore a variance of 1,995 minutes², a number that means nothing to a reader on sight
Standard deviation (SD)Square root of the variance, so back in the original unitsSymmetric, roughly normal data

A standard deviation is only interpretable if the distribution is roughly symmetric, because its usefulness comes entirely from the normal distribution's fixed proportions: about 68 % of observations within one SD of the mean, 95 % within two, and 99.7 % within three (strictly, 95 % falls within 1.96 SD — which is where the 1.96 in every confidence interval comes from). Quote an SD on a heavily skewed variable and "mean ± 2 SD" will include impossible values — negative door-to-CT times, negative lactates. That is the quickest way to spot a summary that should have been a median.

See the normal distribution figure and the skewed-dataset figure, where the same 40 numbers are shown as a histogram, a box plot and a summary table.

SD versus SEM — the confusion worth fixing permanently
SEM = SD / √n

They answer different questions and are not interchangeable.

  • SD describes the patients. It says how spread out the individuals are. It does not shrink as you enrol more of them — a bigger study measures the same underlying variability more precisely, it does not make people more alike.
  • SEM describes the estimate. It says how precisely you have pinned down the mean, and it does shrink with √n. It is the building block of the confidence interval: roughly mean ± 2 SEM.
Because SEM is always smaller than SD, error bars drawn as SEM look tighter. A figure whose bars are SEM tells you about confidence in the average; one whose bars are SD tells you about the spread of patients. They can differ several-fold — in the worked dataset on the Figures tab, SD is 44.7 minutes and SEM is 7.06. If a figure legend does not say which it is plotting, you cannot interpret the bars, and that omission is common enough to be worth noticing every time.
Centiles, and reference ranges

A centile is the value below which that percentage of observations fall; the median is the 50th. Paediatric growth and observation charts are built from them, which is why "below the 2nd centile" is a statement about a population, not a diagnosis.

Most laboratory reference ranges are the central 95 % of a healthy reference population — which has a consequence people rarely state out loud: 1 healthy person in 20 will fall outside the range for any given test, by construction. Order a panel of 20 independent tests on a well person and the expected number of "abnormal" results is one. A reference range is not a threshold for disease, and an isolated marginal abnormality on a broad panel is usually the arithmetic working as designed.

What to report, and therefore what to expect. Symmetric data: mean and SD. Skewed data, ordinal data, or anything with outliers: median and IQR. Paired with the sample size in both cases, because none of these numbers mean anything without an n. The descriptive statistics calculator computes all of them from a series you paste in, and draws the box plot and histogram, so you can see what a summary hides.

You will rarely choose a test yourself. You will constantly need to judge whether the one a paper used was reasonable — and that judgement reduces to three questions asked in order.

The three questions
  1. What kind of outcome is it? Continuous, ordinal, or categorical (see chapter 03).
  2. How many groups are being compared? Two, or more than two. (A single group against a known reference value is the one-sample case, and takes a one-sample t-test, or a Wilcoxon signed-rank test against that value.)
  3. Are the observations paired or independent? Paired means the same patient measured twice, or patients matched one to one — before and after, left eye and right eye, two methods on the same blood sample. Analysing paired data as though it were independent throws away the pairing, which is usually the whole point of the design.
The map
OutcomeTwo independent groupsTwo paired measurementsThree or more groups
Continuous, roughly normalUnpaired (Student's) t-testPaired t-testANOVA
Continuous, skewed; or ordinalMann–Whitney U (Wilcoxon rank-sum)Wilcoxon signed-rankKruskal–Wallis
Categorical / binaryChi-square; Fisher's exact when expected counts are smallMcNemar's testChi-square across the table
Time to eventLog-rank; Cox regression for adjustment—Log-rank across strata

Two practical notes on that table. Fisher's exact is the safe choice whenever any expected cell count is small (the usual rule of thumb is under 5) — it is what the fragility calculator on this site uses. And ANOVA answers only "is there a difference somewhere"; it does not say which groups differ, which requires post-hoc comparisons with an explicit correction for multiplicity.

Parametric or not

"Parametric" tests assume the data follow a particular distribution — usually the normal — and additionally, for the t-test and ANOVA, that group variances are similar and observations are independent. Non-parametric tests replace the values with their ranks and so assume much less.

  • When the assumptions hold, parametric tests have more power, which is the whole reason to prefer them.
  • Skew alone is less dangerous than it is usually made to sound, provided the groups are the same size. The central limit theorem applies to the mean rather than to the data, so with balanced arms a t-test on heavily skewed data holds its false-positive rate at or below the nominal 5 % even at a dozen per group. What it loses is power — on skewed data a t-test can need appreciably more patients than a rank test to detect the same real difference.
  • The genuinely dangerous combination is unequal group sizes together with unequal spread, and it is far more common in published work than skew alone. Student's t-test assumes equal variances; break that and the balance at the same time and its error rate is not approximately anything — it can reject a true null far more often than 5 % of the time when the smaller group is the more variable one, and almost never when it is the less variable one. Welch's correction, which does not assume equal variances, is the default in most software for exactly this reason, and a paper comparing a small group against a much larger one should be using it.
  • A non-parametric test on data that were fine either way costs a little power and almost never misleads. Reaching for one is a mild inefficiency, not an error.
Testing for normality and then choosing the test is not as respectable as it looks. Normality tests are underpowered in exactly the small samples where the assumption matters, and over-powered in the large samples where it does not — they will reject normality for a trivial deviation in 5,000 patients. The judgement should come from the plotted distribution and from what the variable is, not from a p-value about a p-value's assumptions.
Correlation, regression and agreement are three different questions
QuestionToolGives
Do these two move together?Pearson's r (linear, continuous, normal); Spearman's rho (ranks — monotonic, ordinal, robust to outliers)−1 to +1. r² is the proportion of variance shared
Can I predict one from the other, and by how much?Linear regressionA slope in real units, with a confidence interval
Can this method replace that one?Bland–AltmanBias and limits of agreement, in real units
A high correlation does not mean two methods agree. Two measurements can correlate almost perfectly while one reads systematically higher than the other — correlation is unchanged if one method reads a fixed amount high, or a fixed percentage high. A method-comparison paper that offers only an r value has answered the wrong question. The agreement figure shows the same simulated data as a scatter plot with r = 0.990 and as a Bland–Altman plot revealing a bias of about 7 kg that grows with size.
What no test tells you

Every test above returns a p-value, and a p-value is a statement about compatibility with a null hypothesis and nothing else. It does not give the size of the effect, its direction in clinically useful units, or its precision. A paper whose results section is a list of p-values has reported the least informative summary available. What you want, and should look for instead, is the estimate with its confidence interval — a difference in means, a difference in medians, a risk difference — in the units a patient would recognise. See chapter 11.

Bias is systematic error — it moves the estimate in a direction and does not shrink when you enrol more patients. That last clause is the whole reason bias matters more than imprecision: a bigger study fixes imprecision and entrenches bias. The five domains below follow a trial through its own timeline, which is also the order in which to look for them.

DomainWhenWhat happensLook for
Selection biasAt allocationThe groups differ at baseline in ways that affect the outcomeHow the sequence was generated; whether allocation was concealed; the baseline table
Performance biasDuring the trialThe groups get different co-interventions or attentionBlinding of patients and clinicians; protocol for concomitant care; crossover
Detection biasAt measurementOutcomes are sought or scored differently between groupsBlinded outcome assessment; adjudication committee; objectivity of the outcome
Attrition biasAt follow-upLosses differ between groups, or from those who stayThe CONSORT flow diagram; numbers analysed versus randomised; how missing data were handled
Reporting biasAt publicationOutcomes or whole studies are selectively reportedThe registered protocol; whether the primary outcome changed; whether all pre-specified outcomes appear
Three that deserve names of their own
  • Immortal time bias. If a patient must survive long enough to receive the treatment, the treated group is composed of survivors before treatment does anything. This makes ineffective treatments look protective in database studies, and it is subtle enough to reach print regularly.
  • Confounding by indication. Sicker patients get the more aggressive treatment, so the aggressive treatment looks harmful. This is the reason observational comparisons of, say, intubation strategies mislead so reliably.
  • Lead-time and length bias. Screening-detected disease appears to survive longer partly because the clock started earlier, and because slow-growing disease is over-represented among cases found by screening.
The practical test for bias. For each domain, ask: if this went wrong, which way would it push the result, and by enough to matter? A study can be flawed in a direction that works against its own conclusion — unblinded assessment in a trial that reported no benefit makes the null result more, not less, believable. Bias that opposes the finding strengthens it. Naming a flaw is half an appraisal; naming its direction and plausible size is the whole of one.

A confounder is associated with the exposure, is a cause of the outcome independently of the exposure, and is not simply a step on the causal path between them. That third clause matters: adjusting for a mediator removes part of the effect you are trying to measure, and adjusting for a collider — a common consequence of exposure and outcome — can manufacture an association out of nothing.

The tools, and what each actually buys
ToolHandlesResidual problem
RandomisationAll confounders, known and unknownOnly in expectation — small trials can still be unbalanced by chance
Restriction / matchingThe variable you choseLoses generalisability; over-matching can remove real effect
StratificationA few categorical variablesRuns out of patients quickly
Multivariable regressionSeveral measured confoundersNeeds the right variables, correctly measured, with a sensible model form
Propensity scoreMany measured confounders at onceSame dependence on measured variables — it is bookkeeping, not magic
Instrumental variable / Mendelian randomisationUnmeasured confounding, in principleRequires an instrument that is genuinely unrelated to the outcome except through exposure
"Adjusted for" is not "accounted for". Every method except randomisation can only adjust for what was measured, and only as well as it was measured. A propensity model with forty covariates is still blind to illness severity if severity was never recorded — and in emergency medicine, severity is the confounder that matters most and is recorded worst. When an observational study's adjusted estimate is dramatically different from its crude one, that is a warning that confounding dominates the data, not reassurance that it has been removed.

Table 2 fallacy. The other coefficients in a multivariable model are not each a valid causal estimate for their own variable; the model was specified to estimate one exposure effect, and the covariates' coefficients carry a mixture of direct and confounded signal. A paper that presents its whole adjustment table as a list of independent risk factors is over-reading its own model.

These three are routinely conflated, and they protect against different failures at different moments.

  • Sequence generation decides the order of assignments. It must be genuinely random — computer-generated, often in blocks, sometimes stratified by centre or severity. Alternation, date of birth and hospital number are not random and are trivially predictable.
  • Allocation concealment stops the person enrolling a patient from knowing the next assignment. This is the step that prevents selection bias, and it is possible even when blinding is impossible: you can conceal allocation for a surgical trial that nobody can blind. Sealed sequentially numbered opaque envelopes, central telephone or web randomisation, and identical numbered treatment packs all achieve it. CRASH-2 and HALT-IT both used the numbered-pack method: eight identical packs in a box, differing only by number, so the clinician could not steer a sicker patient towards the drug.
  • Blinding (masking) stops people knowing the assignment afterwards, and prevents performance and detection bias. It applies separately to patients, treating clinicians, outcome assessors, data analysts and adjudication committees, and a paper should say which of those were masked rather than claiming to be "double-blind" and leaving you to guess.
Which outcomes need blinding most. The more judgement an outcome requires, the more blinding matters. All-cause mortality is nearly unfalsifiable by an unblinded assessor. "Clinical improvement", pain score, time to discharge, and any composite containing a discretionary component are all highly susceptible. So an unblinded trial with a hard primary outcome may be perfectly credible, while a blinded trial whose primary outcome is a subjective scale still needs care.

Baseline tables and their misuse. The baseline table exists to let you judge whether randomisation produced comparable groups, and to see whom the trial enrolled. It should not carry p-values: in a properly randomised trial any baseline difference is by definition chance, so testing it answers a question nobody asked. What you want instead is the size of the imbalance and whether the imbalanced variable is prognostically important.

Outcome choice is where a trial decides how useful it is allowed to be, and it is decided before a single patient is enrolled.

Surrogate outcomes

A surrogate stands in for what you care about: lactate clearance for survival, radiographic union for function, viral load for illness. Surrogates make trials smaller and faster, and they mislead whenever the intervention affects the marker by a route that does not run through the outcome. The question to ask is not "is this surrogate correlated with the outcome" — it usually is — but "do interventions that move this surrogate reliably move the outcome?" That is a far higher bar, and most surrogates fail it.

Composite outcomes

Composites (death, MI or urgent revascularisation; death or dependency) buy statistical power by collecting more events. They are legitimate when the components are of similar importance to the patient, of similar frequency, and plausibly moved in the same direction by the intervention. They mislead when a frequent, mild, discretionary component — usually something like "unplanned re-attendance" — drives the whole result while the component anyone cares about does not move. Always look for the component-by-component breakdown; if it is absent, treat the composite as unresolved.

The primary outcome is singular. A trial has one primary outcome (or a small pre-specified set with a stated multiplicity plan). Everything else is secondary and exploratory, and is there to generate hypotheses and to make the primary result coherent — not to rescue it. When a trial's abstract leads with a secondary outcome, the primary result was not what the authors hoped. PROPPR is the honest version of this situation: both co-primary mortality outcomes were null, and the paper's conclusion says so in its first clause before reporting the positive secondary findings. See the worked appraisal.
Patient-important outcomes and PROMs

Mortality, function, symptom burden, and days alive and out of hospital are outcomes patients recognise. Patient-reported outcome measures (PROMs) capture them directly, at the cost of needing a validated instrument, a minimal clinically important difference (MCID) established independently of the trial, and an honest account of missing questionnaires. A statistically significant change on a scale is meaningless without knowing the MCID, and quoting one from the same trial that found the effect is circular.

PopulationWho is analysedPreserves randomisationAnswers
Intention-to-treat (ITT)Everyone randomised, in their allocated group, whatever happened nextYesWhat happens if I adopt this policy
Modified ITT (mITT)ITT minus a pre-defined group, e.g. those who never received a dosePartly — depends entirely on the exclusion ruleSomewhere between the two; needs scrutiny
Per-protocolOnly those who completed treatment as specifiedNoWhat happens in patients who tolerate and comply — a self-selected group
As-treatedGrouped by what they actually receivedNoAn observational question inside a trial

Why ITT is the default. Dropping non-compliers destroys the one thing randomisation gave you, because non-compliance is not random: people stop drugs because of side effects, because they deteriorate, or because they recover. Per-protocol analysis therefore compares a filtered intervention group with an unfiltered control group, and the filter is prognostic. ITT preserves the comparison and answers the question a clinician actually faces, which is "if I prescribe this, what happens" — including the patients who will not take it.

ITT is conservative for benefit, not for harm. Diluting an intervention with people who did not receive it pulls an effect estimate towards no difference. So an ITT analysis showing benefit is robust. But the same dilution understates harm, since only the treated can be harmed by treatment — which is why safety is often better assessed on an as-treated basis, and why a trial should report both.

Reading the numbers-analysed line is not pedantry. CRASH-2 randomised 20,211 patients and analysed 10,060 of 10,096 and 10,067 of 10,115 — losses of well under 1 % in both arms, which is why its ITT analysis is credible at face value. HALT-IT states the opposite arrangement plainly: its primary analysis excluded patients who received neither dose and those without outcome data, which is a modified ITT, and the paper says so in its methods rather than leaving you to discover it. A trial that randomises 500 and analyses 380 has a problem regardless of what it calls the analysis.

Missing data
  • Complete-case analysis assumes data are missing completely at random, which they almost never are.
  • Last observation carried forward is not conservative; it can bias either way depending on the disease trajectory.
  • Multiple imputation is the usual defensible approach, and is only as good as the variables predicting missingness.
  • A sensitivity analysis under a plausibly unfavourable assumption is the part that actually reassures. If a conclusion survives assuming every lost patient in the treatment arm did badly, believe it.

Take a trial with a control event rate CER and an experimental event rate EER.

RR = EER / CER   |   RRR = 1 − RR   |   ARR = CER − EER   |   NNT = 1 / ARR

The relative measures answer "how much does this shift a person's risk", and they travel reasonably well between populations of different baseline risk. The absolute measures answer "how many people benefit", and they do not travel at all — they depend entirely on baseline risk. This is why a relative risk reduction quoted without a baseline risk is uninterpretable, and why it is the favourite number of anyone selling something.

Worked: the same RR, two different interventions
SettingCEREERRRRARRNNT
High-risk population20 %15 %25 %5 %20
Low-risk population0.4 %0.3 %25 %0.1 %1000

Identical relative effect; fifty-fold difference in how many people you must treat. If the intervention carries a 1-in-200 risk of serious harm, it is clearly worth doing in the first row and clearly not in the second — and no statistical test will tell you that, because it is a judgement about magnitude, not significance.

Odds ratios

An odds ratio is (a/b)/(c/d) on the 2×2 table. It is the natural output of logistic regression and of case–control studies, where risks cannot be estimated. Its trap is simple: the OR exaggerates the RR, and increasingly so as the outcome becomes more common. For an outcome affecting 2 % of controls, OR and RR are nearly identical; for one affecting 40 %, an OR of 2.0 corresponds to a risk ratio nearer 1.4. Treating a published OR as a risk ratio is one of the most common misreadings in clinical conversation.

Hazard ratios

A hazard ratio compares the instantaneous rate of the event between groups, using time-to-event data, and assumes proportional hazards — that the ratio is roughly constant over follow-up. When survival curves cross or diverge late, that assumption fails and the single HR hides the shape of what happened. Look at the curves and the numbers-at-risk row beneath them; late-separating curves with thin numbers at risk are where over-interpretation lives.

NNT needs four things attached to be meaningful: the time horizon (an NNT of 20 over 30 days is not an NNT of 20 over five years), the baseline risk it was computed at, a confidence interval, and its partner the number needed to harm. An NNT quoted bare is a marketing number. Note also that a confidence interval for ARR that crosses zero produces an NNT interval that passes through infinity and emerges as a number needed to harm — which is mathematically awkward and clinically informative: it means the data are compatible with both help and harm.

The treatment-effect calculator computes all of these with confidence intervals from raw event counts.

What a p-value is

A p-value is the probability of observing data at least as extreme as yours if the null hypothesis were true and the model were correct. That is all it is. It is a statement about data given a hypothesis, not about a hypothesis given data, and the difference is not academic.

p does not meanWhy not
The probability the null is trueThat is an inverse probability and needs a prior. p cannot supply one.
The probability the result is a flukeSame error, more casually phrased.
The size or importance of the effectp mixes effect size with sample size. A trivial effect in a huge trial gives a tiny p.
p > 0.05 means no effectAbsence of evidence. A wide interval around a null point estimate is ignorance, not equivalence.
p = 0.049 and p = 0.051 differ meaningfullyThe threshold is a convention, not a property of nature.
Two trials with p < 0.05 and p > 0.05 disagreeTheir intervals usually overlap heavily. Compare estimates and intervals, not verdicts.
Why the confidence interval is the better tool

A 95 % confidence interval is the range of effects reasonably compatible with the data. It carries everything the p-value carries (if it excludes the null, p < 0.05) and adds the two things a clinician needs: the plausible size of the effect, and the precision of the estimate. Read both of its ends and ask what each would mean clinically:

  • Both ends clinically important, same direction — a useful positive result.
  • One end important, the other trivial — a positive result that has not established a worthwhile effect.
  • Interval spans the null but both ends trivial — genuinely reassuring; this is what "no important difference" looks like.
  • Interval spans the null and both ends important — an uninformative trial, whatever its p-value. ANDROMEDA-SHOCK's 28-day mortality hazard ratio of 0.75 (95 % CI 0.55 to 1.02, p = 0.06) is exactly this: compatible with a 45 % relative reduction in death and with a 2 % increase. Calling that "no difference" discards half the information. See the worked appraisal.
The 95 % interval is not a 95 % probability that the true value lies inside it. Strictly, it is a procedure that captures the true value 95 % of the time across repeated identical experiments. The Bayesian object that does mean what everyone wants it to mean is the credible interval, and it requires a stated prior. In practice the loose reading does little harm at the bedside, but it is worth knowing that it is wrong, and why, for the day someone corrects you.
Multiplicity

Test twenty independent true nulls at α = 0.05 and you expect one "significant" result. Trials handle this by nominating a single primary outcome, pre-specifying a small number of comparisons, and either splitting alpha (Bonferroni, hierarchical testing) or accepting that everything else is exploratory. PROPPR illustrates the honest handling: 23 pre-specified complications were examined and none differed, and the paper reports that as a safety statement rather than mining it for a finding.

The previous chapter ended on an awkward admission: the confidence interval does not mean what almost everyone takes it to mean, and the p-value answers a question nobody asked. Bayesian inference is the framework that answers the question people actually have — given these data, how likely is it that this treatment works? — at the cost of requiring you to state what you believed beforehand.

The engine
prior × likelihood → posterior
  • The prior is what was believed about the effect before this study: a distribution over possible effect sizes, not a single guess.
  • The likelihood is what the new data say, on their own.
  • The posterior combines them, and is the answer — a full distribution over effect sizes, from which any question can be read off directly.
You already do this, every shift. The pre-test to post-test calculator on this site is Bayes' theorem, written in odds. The pre-test probability is the prior. The likelihood ratio is the likelihood. The post-test probability is the posterior. Nobody finds that controversial, or demands to know where the pre-test probability came from before allowing a d-dimer to be interpreted — and the objection most often raised against Bayesian trial analysis is exactly that demand, applied selectively.
Credible intervals versus confidence intervals
95 % confidence interval95 % credible interval
MeansA procedure that captures the true value 95 % of the time across repeated experimentsGiven the data and the prior, a 95 % probability that the true value lies inside
RequiresNo priorA stated prior
Can sayNothing directly about the probability of benefitA statement of the form "there is a 94 % probability that this treatment reduces mortality" — a real worked example follows below

That last row is the practical difference, and it is the reason Bayesian re-analyses of intensive care and emergency trials have become common. A frequentist trial reporting p = 0.06 is obliged to say it failed to reject the null; a Bayesian analysis of the same patients can put a probability on benefit directly.

ANDROMEDA-SHOCK has actually had this done to it, which makes it the cleanest worked example available. The trial reported 28-day mortality with a hazard ratio of 0.75 (95 % CI 0.55 to 1.02, p = 0.06) and concluded, correctly by its own rules, that the strategy did not reduce mortality — see the worked appraisal. A published Bayesian reanalysis of the same patients (Zampieri and colleagues, Am J Respir Crit Care Med 2020) fitted four different priors — optimistic, neutral, null and pessimistic — and found the posterior probability that peripheral-perfusion-targeted resuscitation is superior at 28 days was above 90 % for every one of them. At 90 days it was above 90 % for all but the pessimistic prior. Under the optimistic prior the posterior median odds ratio for 28-day mortality was 0.61 (95 % credible interval 0.41–0.90), against a frequentist odds ratio on the same data of 0.61 (95 % CI 0.38–0.92).

Three things in that are worth more than the headline. The point estimates are identical — the Bayesian analysis did not manufacture an effect, it re-expressed the uncertainty in a form that can carry a probability. The authors did exactly the sensitivity analysis across priors that the section below tells you to look for, rather than reporting one prior and stopping. And the conclusion did shift with the prior at 90 days but not at 28, which is the honest way to show a reader where the data stop deciding and the prior starts.
Where the prior comes from, and why it is not cheating

The honest objection to Bayesian analysis is that the prior can be chosen to get the answer you want. The honest answer is that a good Bayesian analysis makes the prior explicit and then shows how much it mattered. Look for:

  • A vague or reference prior, which deliberately contributes almost nothing, so the posterior is driven by the data alone. With a vague prior the credible interval and the confidence interval are usually near-identical — which tells you the Bayesian machinery is not doing the work.
  • A sceptical prior, centred on no effect and tight enough to make a large claimed benefit implausible. If a result survives a sceptical prior, it is robust.
  • An enthusiastic prior, centred on the hoped-for effect, used as the other bookend.
  • An evidence-based prior from previous trials or a meta-analysis — the most defensible and the least often available.
The sensitivity analysis is the part to read. A Bayesian paper that reports one posterior from one prior has shown you one opinion updated once. A trustworthy one reports the posterior under vague, sceptical and enthusiastic priors side by side, and the reader judges from the spread. If the conclusion flips between a sceptical and an enthusiastic prior, the data are not deciding the question — the prior is, and the paper should say so.
Reading a Bayesian result
  • Posterior probability of benefit — P(effect favours treatment). Beware that this is often near 90 % even for trials everyone agrees were unconvincing; "more likely than not" is a low bar.
  • Probability of a clinically important benefit — P(effect exceeds some threshold you would act on). Far more useful, and it forces someone to name the threshold.
  • Region of practical equivalence (ROPE) — a band around zero within which the effect, if real, would not change practice. Asking how much of the posterior falls inside it is a direct way to say "this difference does not matter", which frequentist analysis cannot do at all.
  • Bayes factor — how much the data shift the odds between two hypotheses, independent of how likely you thought each was to begin with. Note that it is not free of assumptions about effect size: a Bayes factor depends on the distribution of effects assumed under the alternative, so two analyses of identical data can report different Bayes factors because they specified the alternative differently.
When it changes the reading of a trial

Bayesian analysis is most valuable exactly where the frequentist framework is least informative: trials that miss the threshold, trials that stop early (where the framework has no multiplicity penalty, because the posterior does not care how many times you looked), small trials in rare conditions where a well-justified prior adds real information, and adaptive designs. It is not a way of rescuing a negative trial — applied to a result with a tight interval around no effect, it returns a posterior tightly centred on no effect, which is the same conclusion in different words.

Both frameworks describe the same data. They differ in what they are willing to put a probability on, and the practical instruction is the same for both: read the interval and the effect size, not the verdict.

Power is the probability of detecting an effect of a specified size if it is really there. It rises with the number of events (not patients), with the size of the effect sought, and with tolerance for a larger α. Four quantities are locked together — α, power, effect size and sample size — so fixing any three fixes the fourth.

The optimism trap. Sample size is calculated from an assumed effect. Under-funded trials assume implausibly large effects to make the arithmetic affordable, then report "no significant difference" for an effect they were never able to detect. Always find the sample-size paragraph and ask: was the assumed effect plausible? A trial powered for a 15 % absolute mortality reduction in a condition where 3 % would be a triumph has told you almost nothing by failing.
Post-hoc power is a dead end

Calculating power after the fact from the observed effect adds no information — it is a deterministic function of the p-value. If a trial was inconclusive, the confidence interval already says so, and says it better. Report the interval; do not compute observed power.

Interim analyses and early stopping

Repeatedly testing accumulating data inflates the type I error, so trials use alpha-spending (O'Brien–Fleming, Pocock) with a data monitoring committee and pre-specified stopping rules for benefit, harm or futility. Two practical consequences:

  • Trials stopped early for benefit systematically overestimate the effect. You stop at a random high point in the accumulating estimate, and the fewer the events at stopping, the greater the overestimate.
  • Stopping for futility is usually sound, and stopping for harm is mandatory — but neither licenses the confident negative conclusion that often follows.
Fragility

The fragility index is the number of patients whose outcome would have to change, in one arm, to move a significant result across p = 0.05. It is an intuition pump rather than a formal inference: a trial whose headline depends on two events is not the same evidence as one whose headline depends on eighty, even when both report p = 0.04. Its limitations are real — it depends on the test used, ignores effect size and time-to-event data, and can be gamed — so use it to calibrate confidence, not to overturn a result. The fragility calculator computes it with Fisher's exact test.

Subgroup analysis is where most over-interpretation in clinical medicine originates, and also where some of its most important findings came from. The difference between the two is knowable, and it is not whether the subgroup result is "significant".

The wrong way: testing within subgroups

Splitting a trial into subgroups and looking for which ones have p < 0.05 is close to guaranteed to mislead. Each subgroup is smaller and therefore less powered, so significance within subgroups tracks subgroup size as much as effect. The famous demonstration is the ISIS-2 astrological subgroup: aspirin appeared not to work in patients born under Gemini and Libra. Nobody believes that, and the statistics that produced it are the same statistics that produce the subgroup claims people do believe.

The right way: a test for interaction

The question is not "is the effect significant in this subgroup" but "is the effect different between subgroups". That is a formal test for interaction (or, for an ordered variable, for trend), and it should be pre-specified, few in number, biologically motivated, and consistent across related outcomes.

The CRASH-2 timing analysis is the instructive case, in both directions. The 2011 exploratory analysis found strong evidence that the effect of tranexamic acid on death due to bleeding varied by time from injury — test for interaction p < 0.0001. Treatment within one hour reduced death due to bleeding (5.3 % vs 7.7 %, RR 0.68, 95 % CI 0.57–0.82), between one and three hours also reduced it (RR 0.79, 0.64–0.97), and treatment after three hours appeared to increase it (4.4 % vs 3.1 %, RR 1.44, 1.12–1.84). The authors label the analysis exploratory in its own title, and in the same paper report no evidence of variation by blood pressure, GCS or injury type — the discipline of reporting the negative interactions alongside the positive one is what makes the positive one credible. Practice changed on this analysis, and reasonably so: the interaction was strong, monotonic, and mechanistically coherent for an antifibrinolytic, which is meant to act while fibrinolysis is still running. It remains a subgroup finding, and the 3-hour harm signal in particular is a post-hoc estimate that should be quoted with that caveat attached.
A checklist for any subgroup claim
  1. Was it pre-specified, and is the protocol available to check?
  2. How many subgroups were examined in total? (The denominator is rarely stated.)
  3. Is there a formal interaction test, and what is its p-value?
  4. Is the direction biologically plausible, and does it hold for related outcomes?
  5. Is the gradient monotonic where the variable is ordered?
  6. Has it replicated anywhere else?

One or two yeses is a hypothesis. Most of them, as with CRASH-2 timing, is something you can act on while still calling it what it is.

Apply the checklist to CRASH-2 timing honestly and one box stays empty. Questions 3, 4 and 5 are clearly satisfied and question 2 is answerable, because the authors report the null interactions as well as the positive one. But question 1 is not settled by the paper: the analysis is labelled exploratory in its own title, and the published report does not let a reader confirm that the time-to-treatment hypothesis was registered in advance — you would have to go to the protocol. That is exactly why question 1 asks whether the protocol is available to check, rather than whether the authors say it was pre-specified. An honest appraisal names the box it could not tick.

A superiority trial tries to show a difference and treats failure as failure. A non-inferiority trial tries to show the new treatment is not unacceptably worse than the old — usually because it is cheaper, safer, oral, or available at 3 a.m. Its logic is inverted, and so are several of its incentives.

The margin is the whole design

Everything hangs on the pre-specified non-inferiority margin Δ: the largest loss of efficacy you are willing to accept. Set it narrow and the trial is honest and expensive. Set it wide and almost anything passes. The margin should be justified from the size of the comparator's own established effect over placebo — preserving a stated fraction of it — and not from what makes the sample size affordable. A non-inferiority trial with no justification for its margin cannot be appraised, only believed or not.

How to read the interval
Confidence interval versus Δ and the nullConclusion
Entirely on the good side of Δ, crosses the nullNon-inferior, not superior
Entirely on the good side of Δ and excludes the null favourablyNon-inferior and superior
Crosses ΔNon-inferiority not established — inconclusive, not "equivalent"
Entirely beyond ΔInferior
Two inversions that catch people out. First, sloppiness biases towards the desired conclusion: dropouts, non-compliance, crossover and measurement noise all pull results towards no difference, which is what a non-inferiority trial wants to find. So for these designs, per-protocol analysis is presented alongside ITT and both must agree — the reverse of the superiority convention. Second, assay sensitivity: if the comparator underperformed in this trial, non-inferiority may mean only that both arms did badly. Check that the control arm's event rate matches historical expectation.

Equivalence trials are two-sided non-inferiority: the interval must sit inside ±Δ. They are rare in clinical medicine outside bioequivalence. What is common, and wrong, is calling a failed superiority trial "equivalent" — you cannot switch design after the fact, because the failed trial never specified a margin.

The table
Disease presentDisease absent
Test positiveTP (a)FP (b)
Test negativeFN (c)TN (d)
Sensitivity = a/(a+c)  |  Specificity = d/(b+d)  |  PPV = a/(a+b)  |  NPV = d/(c+d)
LR+ = sens / (1 − spec)   |   LR− = (1 − sens) / spec

Sensitivity and specificity are read down the columns — they condition on disease status, and are therefore roughly stable across populations of different prevalence (roughly, not exactly — see spectrum bias below). PPV and NPV are read across the rows, so they depend on prevalence and are not properties of the test at all. A d-dimer's NPV in a low-risk outpatient population and in a critical care unit are different numbers for the same assay. Quoting a PPV without the prevalence it was measured at is meaningless.

Likelihood ratios, and why they are the useful ones

Likelihood ratios convert a pre-test probability into a post-test probability, which is what a clinician is actually doing when ordering a test. They work multiplicatively on odds:

pre-test odds × LR = post-test odds   (odds = p / (1 − p))
LR+Effect on probabilityLR−Effect
> 10Large increase — often conclusive< 0.1Large decrease — often conclusive
5–10Moderate increase0.1–0.2Moderate decrease
2–5Small increase0.2–0.5Small decrease
1–2Negligible0.5–1Negligible
The mnemonics, and their limits. SpPIn: a highly Specific test, when Positive, rules in. SnNOut: a highly Sensitive test, when Negative, rules out. Both are approximations of the likelihood-ratio arithmetic and both fail at extremes of pre-test probability. A test with 99 % sensitivity and 50 % specificity has an LR− of 0.02, which sounds decisive — but applied to a patient whose pre-test probability was 80 %, a negative result still leaves about 7 %, which is not a probability of a dangerous diagnosis that anyone should discharge on. That is precisely the situation in which a clinician is most tempted to be reassured. Do the arithmetic with the pre-test to post-test calculator rather than trusting the mnemonic.
Thresholds, ROC curves and AUC

Everything above assumed the test gives a yes or a no. Most do not — troponin, lactate, d-dimer and every clinical score produce a number, and somebody has to choose where to cut it.

How the curve is built

Take every possible cut-off in turn. At each one, compute sensitivity and 1 − specificity, and plot that pair as a point. Join the points and you have the receiver operating characteristic curve. Each point on it is a cut-off, so the curve is not a property the test has in the abstract — it is the complete menu of trade-offs the test offers.

  • The top-left corner is perfection: everyone with disease detected, nobody without it flagged.
  • The diagonal is a coin toss — a test whose curve lies on it carries no information at all.
  • Moving up and right means loosening the threshold: more sensitivity, more false positives. Moving down and left means tightening it.

The ROC figure shows a curve alongside the two distributions it was built from, which is the thing worth internalising: the curve is a summary of how much those two distributions overlap. Where they separate cleanly the curve bulges towards the corner; where they overlap heavily it collapses onto the diagonal.

AUC, and what the number is worth

The area under the curve (AUC, or c-statistic) summarises discrimination across all thresholds at once, and has a clean interpretation: the probability that a randomly chosen patient with the disease scores higher than a randomly chosen patient without it. 0.5 is chance; 1.0 is perfect. The conventional, and entirely arbitrary, descriptive bands are ~0.7–0.8 acceptable, 0.8–0.9 good, above 0.9 excellent.

Three reasons a high AUC does not mean a useful test. It is threshold-free, so it does not tell you that any usable cut-off exists — a test can have an AUC of 0.85 and still offer no single threshold with acceptable sensitivity and specificity for your decision. It weights all thresholds equally, including the many you would never use, while clinically you care only about the region near your operating point. And it says nothing about calibration — see chapter 17. A test's usefulness lives at its threshold, not in its AUC.
Choosing the cut-off

The Youden index is the most commonly used rule: maximise (sensitivity + specificity − 1), which geometrically is the point on the curve furthest above the diagonal. It is marked on the figure. Its hidden assumption is the one that matters: Youden weights a false negative and a false positive equally.

In emergency medicine they are almost never equal. A missed subarachnoid haemorrhage and an unnecessary CT head are not comparable harms, and no arithmetic derived only from the curve knows that. So:

  • For a rule-out decision, fix the sensitivity you require first — often 98–100 % — and read off whatever specificity that costs. The cut-off is then a clinical decision that the curve merely prices.
  • For a rule-in decision, fix specificity and accept the sensitivity.
  • Beware a cut-off that was chosen in the same dataset used to report its performance. Optimising the threshold and then quoting the accuracy achieved at it is a form of overfitting, and the figures will not reproduce in a validation set.
  • Remember that sensitivity and specificity move in opposite directions along the curve, always. Any paper reporting that a new cut-off improved both should prompt you to check whether the population also changed.
Bias specific to diagnostic studies
  • Spectrum bias. Accuracy measured in florid cases against healthy volunteers is always flattering. It must be measured in consecutive patients in whom the test would actually be used.
  • Partial verification / referral bias. If the reference standard is applied preferentially to test-positives, sensitivity is inflated and specificity deflated.
  • Incorporation bias. If the test result contributes to the reference standard, the test is being compared with itself.
  • Differential verification. Test-positives get CT, test-negatives get a phone call at 30 days. These are not the same reference standard.
  • Indeterminate results excluded. Dropping uninterpretable scans improves the numbers and misrepresents clinical reality.

QUADAS-2 in the Checklists tab is the structured way to walk through these.

Emergency medicine runs on prediction rules, and almost every argument about one comes down to confusing the three stages of its life.

StageWhat is doneWhat it establishesWhat it does not
1. DerivationCandidate predictors collected; the rule is fitted to this datasetThat a combination of variables separates outcomes in these dataAnything about other patients. Performance here is optimistic by construction.
2. ValidationThe fixed rule is applied to new patients — internal (split/bootstrap), temporal, or external (different sites, different country)That performance holds outside the derivation setThat anyone will use it, or that using it helps
3. ImpactThe rule is implemented and patient-level outcomes and resource use are measured, ideally in a cluster RCTThat deploying it changes practice for the better—
The single most useful question about any rule: which stage is this paper? A derivation-only rule has not been shown to work anywhere but on the data that built it. Most published rules never reach external validation, and very few reach an impact study — so the ones that have are qualitatively different evidence, and worth knowing by name. The PECARN paediatric head-injury rules are an unusually complete example: derived and validated in the same prospective cohort of 42,412 children, with the validation-population negative predictive value reported separately by age band — 1176/1176 (100.0 %, 95 % CI 99.7–100.0) under two years and 3798/3800 (99.95 %, 99.81–99.99) at two years and over. See the worked appraisal.
Reading a rule's performance honestly
  • What was the outcome, exactly? "Clinically important" outcomes are defined by the authors, and the definition determines the numbers. PECARN's ciTBI is death from TBI, neurosurgery, intubation for more than 24 h, or admission for at least two nights — a definition that deliberately excludes CT findings that change nothing.
  • Which patients were eligible? A rule validated in GCS 14–15 says nothing about GCS 13. This is the single commonest misapplication of a decision rule at the bedside.
  • Is it a rule-out rule or a rule-in rule? Rule-out tools are built for sensitivity and will have poor specificity by design. Criticising a rule-out tool for over-investigating is criticising it for working.
  • What is the miss rate you are accepting, and does it match the clinical stakes? A 2-in-3800 miss rate is acceptable for a condition where the missed cases are non-operative; it is not automatically acceptable elsewhere.
  • Was it intended to override judgement or to support it? Most are explicitly the latter, and most disputes come from using them as the former.
Calibration is not discrimination, and it is the one that gets dropped. Discrimination (the c-statistic) asks whether higher-risk patients score higher. Calibration asks whether a predicted 5 % risk corresponds to an observed 5 % risk. A model can discriminate beautifully and be badly calibrated in a new population — typically over-predicting risk when transported to a lower-prevalence setting. Since the clinical decision is made against an absolute risk threshold, poor calibration breaks the rule even when the AUC looks excellent. Many validation papers report only the AUC.

A prognostic study asks what happens to this patient over time, not whether an intervention helps. Its central technical problem is censoring: most patients have not had the event when the study ends, and simply excluding them would throw away nearly all the data.

  • Kaplan–Meier estimates the survival curve while handling censoring, on the assumption that censoring is non-informative — that those censored had the same future risk as those who remained. This fails badly when people are lost because they deteriorated.
  • The log-rank test compares whole curves, and is the most powerful test available when hazards are proportional. Its weakness is the mirror image of that strength: it accumulates the difference between observed and expected events across the whole follow-up, so when curves cross, an early disadvantage and a late advantage cancel inside the statistic. Two obviously and importantly different curves can therefore return a comfortably non-significant log-rank p. A non-significant log-rank with crossing curves is a reason to look at the plot and at survival on a fixed date, not a reason to conclude there is no difference.
  • Cox proportional hazards gives an adjusted hazard ratio, and assumes proportional hazards across follow-up. Where that fails, the honest reporting is restricted mean survival time, or a hazard ratio quoted separately by period.
  • Competing risks. If a patient can die of something else first, they can never have the event of interest — and Kaplan–Meier will overstate cumulative incidence. In elderly and critically ill cohorts, competing risks are the norm, and the correct tool is a cumulative incidence function, not 1 minus KM.
Read the numbers-at-risk row. The most over-interpreted feature of any survival figure is a late separation of the curves — which is exactly where the fewest patients remain and the widest uncertainty lives. If the numbers at risk have fallen to single figures, the right-hand third of the plot is decoration. If a paper omits the numbers-at-risk row, it has removed the reader's ability to make that judgement.

For appraising a prognostic model specifically, the relevant reporting standard is TRIPOD, and the risk-of-bias tool is PROBAST; for a prognostic factor study, QUIPS.

A systematic review is a study whose unit of analysis is the study: a pre-registered question, an explicit search, duplicate screening and extraction, formal risk-of-bias assessment, and a stated synthesis plan. A meta-analysis is the optional statistical pooling step. A review can be excellent without pooling, and pooling without the rest is not a systematic review.

Fixed versus random effects

A fixed-effect model assumes one true effect and that differences between studies are sampling noise. A random-effects model assumes the true effect varies between studies and estimates the distribution's mean. Random effects gives wider intervals and weights small studies relatively more — which matters, because small studies are the ones most susceptible to bias. Neither model repairs heterogeneity; they make different assumptions about it.

Heterogeneity
StatisticMeansCaution
I²Proportion of variability due to between-study differences rather than chanceNot an absolute measure; unreliable with few studies; a low I² with wide intervals proves nothing
τ²Absolute between-study variance, on the effect scaleMore informative than I² but rarely quoted
Cochran's Q, pTest of homogeneityBadly underpowered with few studies — a non-significant Q is not evidence of homogeneity
Prediction intervalWhere the effect of a future study would plausibly fallThe most honest single summary of heterogeneity, and the least often reported
Heterogeneity is information, not a nuisance. If trials disagree, the interesting question is why — different populations, doses, timing, comparators, outcome definitions or risk of bias. Pooling across a genuine clinical difference produces an average that describes no real patient. The tranexamic acid literature is the standing example: CRASH-2 in trauma reduced all-cause mortality (RR 0.91, 95 % CI 0.85–0.97) while HALT-IT in gastrointestinal bleeding found no reduction in death due to bleeding (RR 0.99, 95 % CI 0.82–1.18) and more venous thromboembolism (RR 1.85, 1.15–2.98). Pooling them to produce "the effect of tranexamic acid on bleeding" would be an arithmetically valid answer to a clinically meaningless question. See the worked appraisal.
Reading a forest plot in order
  1. The studies — how many, how large, and are they the population you care about?
  2. The individual intervals — do they overlap, and do any sit clearly apart?
  3. The weights — is the pooled estimate one large trial with decoration, or a genuine synthesis?
  4. The diamond — its centre and, more importantly, its width against the line of no effect.
  5. The heterogeneity line — I², τ², and a prediction interval if offered.
  6. The risk-of-bias column — and whether a sensitivity analysis restricted to low-risk studies changes the answer. If removing the high-risk studies removes the effect, the effect belongs to the bias.
Publication bias and funnel plots

Small studies with unremarkable results are less likely to be published, so meta-analyses over-represent positive small studies and overestimate effects. A funnel plot shows effect against precision; asymmetry suggests missing small negative studies. Egger's and Begg's tests formalise it. All of them are underpowered below roughly ten studies, and asymmetry has innocent explanations — genuine small-study effects, differing quality, or a real relationship between dose and population risk. Treat funnel asymmetry as a prompt, not a verdict.

Two special forms
  • Individual patient data (IPD) meta-analysis obtains the raw data, allowing consistent outcome definitions, proper time-to-event analysis and credible subgroup work. It is the strongest form and the rarest.
  • Network meta-analysis compares treatments never trialled head to head, through their common comparators. It adds one critical assumption — transitivity — and requires that the trials be similar enough for indirect comparison to mean anything. Check the network diagram for how much of the conclusion rests on indirect evidence, and look for a formal test of inconsistency.

GRADE separates two judgements that are constantly run together: how certain are we about the effect, and what should we therefore do. They can come apart in both directions, and understanding that is most of what a clinician needs from GRADE.

Certainty of evidence, per outcome

Certainty is rated per outcome, not per study and not per guideline. RCT bodies start high, observational bodies start low, and the rating then moves.

Rate down forRate up for
Risk of bias in the contributing studiesLarge effect size
Inconsistency (unexplained heterogeneity)Dose–response gradient
Indirectness (population, intervention, comparator or outcome)All plausible residual confounding would reduce the observed effect
Imprecision (intervals crossing decision thresholds, few events)—
Publication bias—

The four resulting levels — high, moderate, low, very low — are statements about how likely further research is to change the estimate, not about study design alone.

Strength of recommendation

A strong recommendation ("we recommend") means almost all informed patients would choose this, and it can reasonably become a quality standard. A conditional or weak recommendation ("we suggest") means the right choice depends on values and circumstances, and shared decision-making is the point rather than a courtesy. Strength depends on certainty plus the balance of benefits and harms, the variability of patient values, and resource use.

Strong recommendations on low-certainty evidence are legitimate — and common in emergency medicine, because the situations in which you cannot randomise are often the situations in which doing nothing is clearly worse. A lateral canthotomy for orbital compartment syndrome will never have an RCT. Conversely, high-certainty evidence can carry a weak recommendation when the effect is small and the burden real. If you find yourself arguing that a recommendation must be weak because the evidence is thin, you have collapsed the two axes GRADE exists to separate.
Appraising the guideline itself

The tool is AGREE II, but three questions do most of the work: who wrote it and who paid (conflicts declared and managed, chair independence); is the link from evidence to recommendation traceable (does each recommendation carry its evidence and grade, or does the reader have to take it on trust); and is it current (a named review date, with a statement of what has changed). A guideline recommendation quoted without its number and its grade cannot be checked, which is why both belong in any citation of one.

Everything above assumes the paper is an honest report of what was done. That assumption is usually safe and is not free.

Registration and outcome switching

Prospective registration exists so that the primary outcome cannot be chosen after the results are known. Comparing the registry entry with the publication is a five-minute check that catches the single most consequential form of reporting bias: a primary outcome demoted, a secondary promoted, or a new outcome appearing. Registration numbers are given precisely so this is possible — CRASH-2 lists ISRCTN86750102 and NCT00375258; HALT-IT lists ISRCTN11225767 and NCT01658124; PROPPR lists NCT01545232; ANDROMEDA-SHOCK lists NCT03078712.

Spin

Spin is presentation that implies more than the data support without stating anything false. Its recurring forms:

  • A conclusion resting on a secondary or subgroup outcome while the primary was null.
  • "A trend towards benefit" for a non-significant result — there is no such thing as a trend in a single trial; there is an estimate and an interval.
  • Relative risk reductions in the abstract, absolute risks nowhere.
  • "Well tolerated" in place of the harms table.
  • A title that states a causal claim an observational design cannot support.
  • An abstract conclusion broader than the population enrolled.
Read the abstract last. The abstract is the most heavily optimised part of a paper and the only part most readers see. If you read methods, then the results tables, then form a view, then read the abstract, you will notice the gap. If you read the abstract first, you will spend the rest of the paper confirming it.
Conflicts and funding

A declared commercial interest does not invalidate a study, and its absence does not sanctify one — academic, intellectual and career conflicts are real and rarely declared. What matters is whether the funder had a structural role in the parts that can be steered: design, conduct, analysis, and the right to publish a negative result. Note that CRASH-2's funding included both public (UK NIHR Health Technology Assessment programme) and commercial (Pfizer) sources, alongside charitable funders, and that it reported a modest effect on a hard outcome — the combination is worth noting rather than either dismissing or ignoring.

Replication and effect decay

Early positive results are, on average, exaggerated — a consequence of publication bias, small samples, flexible analysis and the winner's curse of reporting whichever estimate crossed the threshold. The practical implication is a prior, not a cynicism: expect a first striking result to shrink on replication, and be slower to adopt on a single trial than the trial's own confidence interval suggests. In a field where a single-centre trial can change practice within a month, that prior is the most useful thing on this page.

And the newer failure modes

Paper mills, fabricated datasets and undisclosed generative-AI text are now part of the landscape, most heavily in low-barrier journals. Signals worth noticing: implausibly clean baseline tables, distributions too uniform, an author list with no traceable institutional link to the work, and reference lists containing citations that do not resolve. Check that a citation exists and says what it is claimed to say — the most reliably detectable defect in an unsound paper is not its statistics but its references.

Qualitative research

Qualitative work answers questions quantitative work cannot: why a pathway is not followed, what a diagnosis means to a patient, how a team actually makes a decision under pressure. Appraising it against RCT criteria is a category error — sample size, blinding and generalisability are the wrong questions. The right ones are:

  • Is the method matched to the question? Grounded theory to build explanation, phenomenology for lived experience, ethnography for practice as it happens, framework analysis for applied policy questions.
  • Sampling — purposive and justified, rather than convenient, and pursued to the point where new data stop adding categories.
  • Reflexivity — has the researcher's own position and its influence on the data been examined and stated?
  • Analytic rigour — independent coding, an audit trail from data to theme, and negative cases sought and reported rather than smoothed away.
  • Are the data shown? Themes should be evidenced with quotations that a reader can weigh, not merely asserted.
  • Transferability, not generalisability — enough contextual description for you to judge whether it applies to your department.

The reporting standards are COREQ (interviews and focus groups) and SRQR; CASP publishes a qualitative checklist, and GRADE-CERQual rates confidence in findings from qualitative evidence synthesis.

Mixed methods

The question to ask is why the two strands were combined and how they were integrated. A paper that reports a survey and an interview study side by side, with no integration, is two papers sharing a title. Genuine mixed-methods work states its sequence (explanatory, exploratory or convergent) and shows where the strands informed each other.

Health economics

An economic evaluation compares costs and consequences of alternatives. Key features to check: the perspective (health service versus societal — it changes which costs count), the time horizon and discount rate, whether the outcome is a cost-effectiveness ratio (cost per event avoided) or a cost–utility ratio (cost per QALY), the incremental cost-effectiveness ratio rather than average costs, and the sensitivity analysis — a probabilistic one with a cost-effectiveness acceptability curve, since a point ICER without uncertainty is not a result. The reporting standard is CHEERS. Be alert to models whose conclusion is driven by one poorly evidenced input; the sensitivity analysis is where that shows, which is why it is the section to read first.

This is the sequence to use on a paper someone has just sent you claiming it changes everything. It is deliberately ordered so that the cheapest disqualifying questions come first.

Ten minutes
  1. Write your PICO question (30 s). What would you change if this were true?
  2. Title, journal, date, design (30 s). Do not read the abstract conclusion yet.
  3. Population (1 min). Eligibility criteria and the baseline table. Is your patient in this trial? This step disqualifies more papers than every statistical consideration combined.
  4. Primary outcome, as pre-specified (1 min). Is it patient-important? Is it a composite? Check the registry entry if something feels moved.
  5. Allocation and blinding (1 min). Concealed? Who was masked? How subjective is the outcome?
  6. The flow diagram (1 min). Randomised versus analysed, in both arms. What happened to the difference?
  7. The primary result, as an absolute number with its interval (2 min). Convert it yourself. Compute the NNT. Read both ends of the interval as clinical scenarios.
  8. Harms (1 min). Was the trial capable of detecting the harm that would matter, and is there a number needed to harm?
  9. Applicability (1 min). Your department's baseline risk, your case mix, whether you can actually deliver the intervention in the time window the trial used.
  10. Now read the abstract's conclusion (30 s), and notice any gap between it and what you just concluded. That gap is the appraisal.
The one-sentence output. A finished appraisal fits in a sentence of this shape: "In [population], [intervention] versus [comparator] changed [outcome] from X % to Y % (absolute difference Z, 95 % CI a to b; NNT n), on a design limited chiefly by [the one dominant flaw], which for our patients means [what you will do]." If you cannot fill every slot, you have not finished reading — and the slot you cannot fill tells you which section to go back to.
Running a journal club that works
  • Pick papers that could change practice, not papers that are easy to criticise. Demolishing a bad paper teaches nothing anyone can use on shift.
  • Circulate the paper and the PICO question in advance, and ask everyone to write down their prediction of the result. Predictions expose priors, and priors are what actually get updated.
  • Assign the domains — one person on population, one on outcome definition, one on the analysis, one on harms. Structure beats enthusiasm.
  • Appoint someone to defend the paper. Journal clubs drift into unanimous criticism, which is comfortable and uninstructive.
  • Finish with a decision, recorded: adopt, adopt with caveats, await replication, or reject — and who will check whether local practice actually changed. A journal club without an owned outcome is a reading group.
  • Keep the appraisals. A department that can find last year's appraisal of the trial now being cited at it has an institutional memory, which is rarer and more valuable than any individual's appraisal skill.
The disposition that matters most

Appraisal is not a search for reasons to dismiss. The commonest failure among people who have just learned it is reflexive scepticism, which is indistinguishable in its effects from credulity: both end in practice unchanged by evidence. The aim is calibration — believing large, well-conducted, replicated findings about patients like yours, believing them proportionately less as each of those conditions weakens, and being able to say out loud which condition is doing the work in any given case.

The arithmetic, done properly

Six calculators covering the numbers appraisal actually requires. All computation happens in your browser; nothing is transmitted or stored. Confidence intervals are labelled with the method used, because a proportion near 0 or 100 % gives nonsense under the textbook normal approximation and these use Wilson's method instead.

01 · Diagnostic accuracy from a 2×2 table
Enter the four cells as counts. Sensitivity, specificity and predictive values come with Wilson 95 % intervals; likelihood ratios with log-method intervals. Remember that the predictive values below apply only at the prevalence of this table — to move a test to your own population, use the pre-test to post-test calculator beneath.
Disease +
Disease −
Test +
Test −
02 · Pre-test to post-test probability
The step a clinician is actually performing when ordering a test. Give a pre-test probability and either the likelihood ratios directly or the test's sensitivity and specificity. The result is the probability of disease after a positive and after a negative result, in your population rather than the study's.
03 · Treatment effect, NNT and confidence intervals
Enter events and totals for each arm. Gives risk ratio (log method), odds ratio (Woolf), absolute risk reduction (Wald) and number needed to treat with its interval. Where the ARR interval crosses zero, the NNT interval passes through infinity and the calculator says so rather than printing a misleading pair of numbers.
Loaded with the CRASH-2 all-cause mortality figures (1463/10060 versus 1613/10067) so you can check the calculator against a published result: the trial reported RR 0.91, 95 % CI 0.85–0.97.
04 · Fragility index
How many patients' outcomes would have to change, in one arm, to move the result across p = 0.05 on a two-sided Fisher exact test. For a significant result this is the fragility index; for a non-significant one the calculator reports the reverse fragility index instead. Treat it as a calibration aid: it ignores effect size, time-to-event structure and the trial's actual analysis, and a large trial can be fragile while a small one is not.
Loaded with the ANDROMEDA-SHOCK 28-day mortality counts (74/212 versus 92/212), a result the trial reported as p = 0.06 on its own time-to-event analysis.
05 · Sample size for two proportions
The calculation whose assumptions you should check in every trial you read. Enter the expected control event rate and the rate you hope to achieve; the result is the number needed per arm, uncorrected for dropout. Try shrinking the effect you are looking for and watch the sample size climb — that curve is the reason under-powered trials assume implausible effects.
06 · Descriptive statistics from a series
Paste or type a list of numbers — separated by commas, spaces or newlines — and this returns everything a results table's first column should contain, with a live histogram and box plot. It is loaded with the 40 simulated door-to-CT times used in the Figures tab. Change one value to an extreme one and watch what happens to the mean and the SD while the median and IQR barely move — that is robustness, and it is the whole argument for reporting the median on skewed data.

Checklists, and what each kind is for

Three different families of tool get used interchangeably and should not be. Reporting guidelines (CONSORT, STROBE, PRISMA, STARD, TRIPOD, CHEERS) tell authors what to write down; a paper that omits an item may still be a good study, badly reported. Risk-of-bias tools (RoB 2, ROBINS-I, QUADAS-2, PROBAST, Newcastle–Ottawa) judge whether the study's result is likely to be distorted. Appraisal checklists (CASP) are teaching scaffolds for a reader forming a judgement.

The questions below are the working substance of each — enough to appraise with, phrased for the reader rather than the author. They are summaries for teaching, not reproductions: for a formal assessment, use the current official instrument from its own publisher.

Is it valid?
  1. Was there a clearly focused question — population, intervention, comparator, outcome?If you cannot restate it in one sentence, neither could the authors.
  2. Was assignment randomised, by a method that is actually random?Alternation, birth date and hospital number are not.
  3. Was allocation concealed from the person enrolling patients?The step that prevents selection bias, and possible even when blinding is not.
  4. Were patients, clinicians and outcome assessors blinded — and which of them?"Double-blind" without saying who is an unanswered question.
  5. Were the groups similar at baseline in prognostically important respects?Judge the size of any imbalance, not its p-value.
  6. Apart from the intervention, were the groups treated equally?Differential co-intervention, monitoring or crossover.
  7. Was follow-up complete, and were losses balanced and explained?Compare randomised with analysed, arm by arm, on the flow diagram.
  8. Were patients analysed in the groups to which they were randomised?ITT. If a modified population was used, is the exclusion rule pre-defined and prognostically neutral?
  9. Was the primary outcome pre-specified, patient-important and objectively measured?Check the registry. A promoted secondary outcome is the commonest form of outcome switching.
  10. Was the trial stopped early, and if so on what pre-specified rule?Stopping for benefit overestimates effect size.
What are the results?
  1. How large is the effect, as an absolute difference with its confidence interval?
  2. What is the NNT, at what baseline risk, over what time horizon?
  3. Read both ends of the interval as clinical scenarios — would each change your practice?
  4. Were harms sought actively, and was the trial large enough to find the ones that matter?
Will it help my patients?
  1. Would my patient have been eligible — and if not, is the difference likely to change the effect or only its baseline risk?
  2. Can I deliver the intervention as trialled, in the time window used, with the same monitoring?
  3. Is my department's baseline risk similar enough for the absolute benefit to carry over?
  4. Do the benefits outweigh the harms and costs for this patient's values?
The formal risk-of-bias instrument for randomised trials is RoB 2, which rates five domains — randomisation process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result — per outcome rather than per study. The reporting guideline is CONSORT, with extensions for cluster, non-inferiority, pragmatic and pilot trials, and SPIRIT for protocols.
Cohort studies
  1. Was the cohort recruited in an acceptable way, and is it representative of the population it describes?
  2. Was exposure measured accurately, and could misclassification differ between those who did and did not develop the outcome?
  3. Was the outcome measured accurately, and by someone unaware of exposure status?
  4. Have the authors identified the important confounders — and is illness severity among them?In acute care, severity is the confounder that matters most and is recorded worst.
  5. Have those confounders been accounted for, and does the adjusted estimate differ dramatically from the crude one?A large shift on adjustment signals that confounding dominates the data.
  6. Was follow-up long enough and complete enough for the outcome studied?
  7. Could reverse causality, immortal time or confounding by indication explain the finding?
  8. Is there a dose–response gradient, and does the effect size survive a negative-control analysis if one was done?
Case–control studies
  1. Were cases defined and ascertained in a way independent of exposure?
  2. Were controls drawn from the same population, and would they have been identified as cases had they developed the outcome?
  3. Was exposure ascertained identically for cases and controls, ideally blind to case status?Recall bias is the design's characteristic flaw; records beat interviews.
  4. Are the results reported as odds ratios, and are they being read as such rather than as risk ratios?
The reporting guideline is STROBE (with RECORD for routinely collected data). The risk-of-bias tool for a non-randomised study of an intervention is ROBINS-I, which asks you to specify the target randomised trial the study is emulating — a discipline worth borrowing even informally, because it forces the confounding question into the open. Newcastle–Ottawa is still widely used inside meta-analyses.

QUADAS-2 assesses four domains for risk of bias, and the first three also for applicability to your own question.

1. Patient selection
  1. Was a consecutive or random sample of eligible patients enrolled, rather than a convenient one?
  2. Was a case–control design avoided?Known cases versus healthy controls inflates accuracy — spectrum bias.
  3. Were inappropriate exclusions avoided?Excluding difficult or indeterminate patients is the commonest way accuracy is inflated.
  4. Applicability: do these patients match the ones in whom you would use the test?
2. Index test
  1. Was the index test interpreted without knowledge of the reference standard?
  2. If a threshold was used, was it pre-specified?A threshold chosen after seeing the data is the diagnostic equivalent of outcome switching, and its reported accuracy will not reproduce.
  3. Applicability: was the test conducted and interpreted as it would be in your department, by operators of comparable experience?
3. Reference standard
  1. Is the reference standard likely to classify the target condition correctly?
  2. Was it interpreted without knowledge of the index test result?Otherwise incorporation bias — the test is partly being compared with itself.
  3. Was the same reference standard applied to everyone?CT for test-positives and a phone call for test-negatives is differential verification.
  4. Applicability: does the reference standard define the condition as you mean it clinically?
4. Flow and timing
  1. Was the interval between index test and reference standard short enough that the condition could not change?
  2. Did all enrolled patients receive both the index test and the reference standard?Partial verification inflates sensitivity.
  3. Were all patients included in the analysis, with indeterminate results reported rather than dropped?
The reporting guideline is STARD. For the review level, the accuracy meta-analysis of a diagnostic test uses a hierarchical or bivariate model rather than pooling sensitivity and specificity separately — because the two are correlated through the threshold, and averaging them independently produces a summary point that no real threshold delivers.
  1. Was the question focused, and was the review registered or protocolised in advance (PROSPERO)?Compare the protocol's primary outcome with the published one.
  2. Was the search comprehensive — several databases, trial registries, reference lists, no arbitrary language or date limits, grey literature considered?
  3. Were screening and data extraction done in duplicate, with disagreement resolved by a stated process?
  4. Were inclusion criteria applied as written, and are excluded studies listed with reasons?
  5. Was risk of bias assessed with an appropriate tool, and used in the analysis rather than merely tabulated?The test that matters: does a sensitivity analysis restricted to low-risk studies preserve the effect?
  6. Were the studies similar enough in population, intervention, comparator and outcome to pool at all?This is a clinical judgement, made before looking at I².
  7. Was heterogeneity quantified and, more importantly, explained — and is a prediction interval given?
  8. Was the choice of fixed or random effects stated and justified?
  9. Was publication bias assessed, and were there enough studies for that assessment to mean anything?Funnel-plot methods are underpowered below roughly ten studies.
  10. Is the pooled estimate presented in absolute terms at a stated baseline risk, not only as a relative measure?
  11. Is the certainty of evidence rated per outcome, ideally with a GRADE summary-of-findings table?
  12. Does the conclusion follow from the whole body of evidence, or from its most striking constituent?
The reporting guideline is PRISMA (with extensions for network meta-analysis, IPD, scoping reviews and abstracts). The critical appraisal instrument for a review is AMSTAR 2, which designates seven of its items as critical — a review failing any of those cannot be rated highly however well it scores elsewhere. ROBIS is the alternative risk-of-bias-in-reviews tool.
The guideline (AGREE II domains, condensed)
  1. Scope and purpose — is the question, population and intended user stated explicitly?
  2. Stakeholder involvement — was the panel multidisciplinary, and were patients involved?
  3. Rigour of development — a systematic search, explicit selection criteria, stated methods for formulating recommendations, external review, and a procedure for updating.The single most discriminating domain, and the one most often thin.
  4. Clarity of presentation — are recommendations specific, unambiguous, numbered, and are options presented?
  5. Applicability — are facilitators, barriers, resource implications and audit criteria addressed?
  6. Editorial independence — is the funder's role stated, and are panel conflicts declared and managed?
Three questions that do most of the work
  • Can I trace each recommendation to its evidence and its grade? If not, the guideline is asking to be trusted rather than read.
  • Is it current, with a named review date? A withdrawn or superseded document circulating as a laminate is a recurring hazard, and the version on a departmental wall is not evidence that it is live.
  • Who paid, and who chaired?
Reading a GRADE summary-of-findings table
ColumnRead it as
Participants / studiesHow much evidence, and of what design
Relative effectThe transportable number
Anticipated absolute effectsThe clinically meaningful number — check which baseline risk it assumes, and whether it is yours
CertaintyHigh / moderate / low / very low, per outcome, with footnotes naming the reason for every downgrade
CommentsWhere the footnotes explaining the downgrades actually live — read these, not the letter grade
Do not collapse certainty into strength. "Strong recommendation, low-certainty evidence" is a coherent and common position — it says the panel is confident about what to do and not confident about the size of the effect, which is the normal state of affairs for time-critical interventions that cannot be randomised. Quoting a recommendation without its grade, or a grade without its recommendation number, removes the reader's ability to check either.

Six papers, six traps

Each of these is appraised for the method lesson it teaches, not as clinical guidance — several are chosen precisely because the clinical message is more complicated than the headline. Every figure, interval and p-value below was taken from the paper's own abstract retrieved from PubMed, and every author line and journal citation verified the same way. Where a number is not in the source, it is not here.

Read each card's lesson box last. The point of the sequence is that six real trials cover most of what goes wrong in reading: intention-to-treat, subgroups and interaction, a null primary with positive secondaries, external validity and harm, derivation versus validation, and the interval that a p-value hides.

01 · Intention-to-treat and the large simple trial
Effects of tranexamic acid on death, vascular occlusive events, and blood transfusion in trauma patients with significant haemorrhage (CRASH-2)
CRASH-2 trial collaborators; Shakur H, Roberts I, Bautista R, Caballero J, Coats T, Dewan Y, El-Sayed H, Gogichaishvili T, Gupta S, Herrera J, Hunt B, Iribhogbe P, Izurieta M, Khamis H, Komolafe E, Marrero MA, Mejía-Mantilla J, Miranda J, Morales C, Olaomi O, Olldashi F, Perel P, Peto R, Ramana PV, Ravi RR, Yutthakasemsunt S. Lancet 2010;376(9734):23–32. PMID 20554319.
Design
RCT Randomised, placebo-controlled, 274 hospitals in 40 countries. Randomisation balanced by centre, block size of eight, computer-generated sequence; participants and study staff (site investigators and trial coordinating centre staff) masked. The identical-numbered-pack mechanism described below is set out in the trial's 2011 analysis, card 02.
Population
20,211 adult trauma patients with, or at risk of, significant bleeding, randomised within 8 h of injury.
Intervention
Tranexamic acid 1 g over 10 min, then 1 g over 8 h, versus matching placebo.
Primary
Death in hospital within 4 weeks of injury, reported by cause category. All analyses by intention to treat.
Result
All-cause mortality 1463 (14.5 %) vs 1613 (16.0 %); RR 0.91, 95 % CI 0.85–0.97; p = 0.0035. Death due to bleeding 489 (4.9 %) vs 574 (5.7 %); RR 0.85, 95 % CI 0.76–0.96; p = 0.0077.
Analysed
10,096 allocated tranexamic acid and 10,115 placebo, of whom 10,060 and 10,067 respectively were analysed.
Funding
UK NIHR Health Technology Assessment programme, Pfizer, BUPA Foundation, and J P Moulton Charitable Foundation. Registered ISRCTN86750102, NCT00375258, DOH-27-0607-1919.

What to notice in the methods. This is the archetype of the large simple trial: a broad eligibility criterion resting on clinician uncertainty, a trivially deliverable intervention, one hard outcome, and enough patients that a small true effect is detectable. The allocation mechanism deserves attention because it achieves concealment without any infrastructure — eight identical numbered packs in a box, so the enrolling clinician cannot know what the next patient will receive and therefore cannot steer sicker patients towards the drug.

Read the absolute numbers. The relative risk of 0.91 sounds modest; the absolute reduction is 16.0 % to 14.5 %, which is 1.5 percentage points and an NNT of about 68 to prevent one death, in a condition that kills one patient in six. Feed the counts into the treatment-effect calculator — it is preloaded with them — and confirm the published interval reproduces.

Where the interval sits. The upper bound is 0.97. That is a positive result whose weakest plausible version is a 3 % relative reduction in death — small, but on an outcome and at a baseline risk where small still matters. This is the pattern described in chapter 11 as "both ends clinically important, same direction", and it is why a modest p-value on a large trial with a hard outcome is stronger evidence than a striking p-value on a small one.

LessonITT is credible when the flow diagram is. "Analysed by intention to treat" is a claim to be checked, not accepted. Here the numbers analysed differ from the numbers randomised by 36 and 48 patients out of ten thousand in each arm — losses small enough that no plausible assumption about them changes the conclusion. The same words on a trial that randomised 500 and analysed 380 would mean something entirely different. Always read the two numbers, arm by arm.
02 · Subgroups, interaction tests, and a post-hoc harm signal
The importance of early treatment with tranexamic acid in bleeding trauma patients: an exploratory analysis of the CRASH-2 randomised controlled trial
CRASH-2 collaborators; Roberts I, Shakur H, Afolabi A, Brohi K, Coats T, Dewan Y, Gando S, Guyatt G, Hunt BJ, Morales C, Perel P, Prieto-Merino D, Woolley T. Lancet 2011;377(9771):1096–1101. PMID 21439633.
Design
Exploratory analysis Secondary analysis of the CRASH-2 randomised population. The word "exploratory" is in the paper's own title.
Question
Whether the effect on death due to bleeding varies by time to treatment, systolic blood pressure, Glasgow coma score, or type of injury. All analyses by intention to treat.
Result
1063 deaths (35 %) were due to bleeding. Strong evidence that the effect varied with time from injury: test for interaction p < 0.0001.
≤1 h: 198/3747 (5.3 %) vs 286/3704 (7.7 %); RR 0.68, 95 % CI 0.57–0.82; p < 0.0001.
1–3 h: 147/3037 (4.8 %) vs 184/2996 (6.1 %); RR 0.79, 0.64–0.97; p = 0.03.
>3 h: 144/3272 (4.4 %) vs 103/3362 (3.1 %); RR 1.44, 1.12–1.84; p = 0.004.
Negative
No evidence that the effect varied by systolic blood pressure, Glasgow coma score, or type of injury.
Conclusion
Tranexamic acid should be given as early as possible; for patients admitted late after injury it is less effective and could be harmful.

Why this subgroup analysis is believable when most are not. Apply the six-question test from chapter 14. The hypothesis was motivated in advance by the drug's mechanism — an antifibrinolytic given after fibrinolysis has resolved has no substrate to act on. The analysis is a formal test for interaction, not a hunt for significance within strata, and the interaction p-value is extreme. The gradient is monotonic across three ordered time bands. And critically, the authors report the interactions that were null in the same breath — blood pressure, GCS and injury type — which tells you the denominator of comparisons examined instead of leaving you to guess it.

Why it is nonetheless a subgroup finding. Randomisation guaranteed comparability within the trial as a whole; it did not guarantee it within a time stratum, because time to treatment is a post-randomisation characteristic correlated with injury pattern, transport distance and survival to enrolment. Patients treated after three hours are a different population from those treated within one, and not only in when they got the drug. The 3-hour signal is therefore the weakest of the three estimates despite its tidy confidence interval.

LessonThe right question is "is the effect different between subgroups", not "is it significant in this one". A pre-specified, mechanistically motivated, monotonic interaction with p < 0.0001, reported alongside its null companions, is about as strong as subgroup evidence gets — strong enough that it changed practice worldwide, and reasonably so. It remains an exploratory estimate, and the apparent harm after three hours in particular should always be quoted with that label attached. Holding both of those thoughts at once is the skill.
03 · A null primary outcome with positive secondaries
Transfusion of plasma, platelets, and red blood cells in a 1:1:1 vs a 1:1:2 ratio and mortality in patients with severe trauma: the PROPPR randomized clinical trial
Holcomb JB, Tilley BC, Baraniuk S, Fox EE, Wade CE, Podbielski JM, del Junco DJ, Brasel KJ, Bulger EM, Callcut RA, Cohen MJ, Cotton BA, Fabian TC, Inaba K, Kerby JD, Muskat P, O'Keeffe T, Rizoli S, Robinson BR, Scalea TM, Schreiber MA, Stein DM, Weinberg JA, Callum JL, Hess JR, Matijevic N, Miller CN, Pittet JF, Hoyt DB, Pearson GD, Leroux B, van Belle G; PROPPR Study Group. JAMA 2015;313(5):471–82. PMID 25647203.
Design
RCT Pragmatic, phase 3, multisite randomised clinical trial at 12 level I trauma centres in North America, August 2012 to December 2013. Local standard-of-care interventions were uncontrolled.
Population
680 severely injured patients arriving directly from scene and predicted to require massive transfusion; 338 allocated 1:1:1 and 342 allocated 1:1:2.
Primary
Co-primary: 24-hour and 30-day all-cause mortality.
Primary result
24 h: 12.7 % vs 17.0 %; difference −4.2 % (95 % CI −9.6 % to 1.1 %); p = 0.12. 30 days: 22.4 % vs 26.1 %; difference −3.7 % (95 % CI −10.2 % to 2.7 %); p = 0.26. Neither was significant.
Secondary
Exsanguination, the predominant cause of death in the first 24 h, was lower in the 1:1:1 group: 9.2 % vs 14.6 %; difference −5.4 % (95 % CI −10.4 % to −0.5 %); p = 0.03. More patients achieved haemostasis: 86 % vs 78 %; p = 0.006.
Safety
The 1:1:1 group received more plasma (median 7 vs 5 U, p < 0.001) and platelets (12 vs 6 U, p < 0.001) and similar red cells (9 U). No differences across 23 pre-specified complications, including ARDS, multiple organ failure, venous thromboembolism, sepsis and transfusion-related complications.

The appraisal problem this poses. A trial that misses both co-primary outcomes and hits two secondaries is the commonest genuinely difficult situation in clinical appraisal, and the temptation runs both ways: to dismiss the trial as negative, or to promote the secondary outcome as the real finding. Neither is right. What the trial licenses is a statement with the hierarchy preserved: 1:1:1 did not reduce mortality at 24 h or 30 days; it did reduce death from exsanguination and improve haemostasis, without a detectable safety cost from the extra plasma and platelets.

Three things make the secondary findings more than noise. They were pre-specified ancillary outcomes, not discovered afterwards. They are mechanistically coherent with the intervention and with each other — a ratio intended to correct coagulopathy reduced deaths from bleeding and produced more haemostasis. And they are consistent in direction with the null primary results, both of which favoured 1:1:1 numerically. A secondary outcome that pointed the opposite way to the primary would deserve far more suspicion.

Read the primary intervals rather than their p-values. The 24-hour difference of −4.2 % has an interval running from −9.6 % to +1.1 %. That is compatible with a large clinically decisive benefit and with a trivial harm — the "uninformative trial" pattern from chapter 11, and a much more honest description than "no difference". The trial is not evidence that ratio does not matter; it is evidence that this trial could not settle it at this sample size.

LessonKeep the outcome hierarchy when you report a trial, including to yourself. The primary outcome is what the trial was designed and powered to answer, and a null primary stays null however attractive the secondaries. The correct response is neither to discard the secondary findings nor to promote them — it is to describe them as what they are, and to notice that the 23-complication safety analysis is doing quiet, unglamorous work here: it is the reason a strategy that was not shown to save lives could still be adopted without evident cost.
04 · External validity, and a harm signal in a null trial
Effects of a high-dose 24-h infusion of tranexamic acid on death and thromboembolic events in patients with acute gastrointestinal bleeding (HALT-IT)
HALT-IT Trial Collaborators. Lancet 2020;395(10241):1927–1936. PMID 32563378.
Design
RCT International, multicentre, randomised, placebo-controlled, 164 hospitals in 15 countries, July 2013 to June 2019. Numbered identical packs; patients, caregivers and outcome assessors masked.
Population
12,009 patients with significant — defined as at risk of bleeding to death — upper or lower gastrointestinal bleeding, enrolled where the responsible clinician was uncertain whether to use tranexamic acid. 5994 tranexamic acid, 6015 placebo; 11,952 (99.5 %) received the first dose.
Intervention
1 g loading dose over 10 min, then 3 g over 24 h at 125 mg/h — a substantially higher total dose over a longer period than the trauma regimen.
Primary
Death due to bleeding within 5 days of randomisation. The analysis excluded patients who received neither dose and those without outcome data on death — a modified ITT, stated in the methods.
Result
222/5956 (4 %) vs 226/5981 (4 %); RR 0.99, 95 % CI 0.82–1.18. Arterial thromboembolic events were similar: 42/5952 (0.7 %) vs 46/5977 (0.8 %); RR 0.92, 0.60–1.39. Venous thromboembolic events were higher with tranexamic acid: 48/5952 (0.8 %) vs 26/5977 (0.4 %); RR 1.85, 95 % CI 1.15–2.98.
Conclusion
Tranexamic acid did not reduce death from gastrointestinal bleeding, and should not be used for it outside a randomised trial. Registered ISRCTN11225767, NCT01658124. Funded by the UK NIHR Health Technology Assessment Programme.

The trial's own premise is the lesson. The background section states that meta-analyses of small trials suggested tranexamic acid might decrease deaths from gastrointestinal bleeding — and a definitive trial of 12,009 patients found a risk ratio of 0.99. That is the small-study effect from chapter 19 playing out in real time, and it is the single best argument on this page for why a pooled estimate from small trials is a hypothesis rather than a conclusion.

Read this null result properly. The interval is 0.82 to 1.18 — tight, centred on no effect, and with both ends clinically unimportant. This is the pattern that genuinely licenses the phrase "no important difference", and it is quite different from ANDROMEDA-SHOCK's 0.55 to 1.02. A null result is only as informative as its interval is narrow, and here it is narrow because the trial was very large.

Why this is not a contradiction of CRASH-2. Different population (gastrointestinal bleeding, not trauma), different dose and duration (4 g over 24 h versus 2 g over 8), different primary outcome and different time frame. Fibrinolysis is central to the coagulopathy of major trauma in a way it is not to variceal or ulcer bleeding. The two trials are not in conflict; they are evidence that "does tranexamic acid help bleeding" is not a single question. Pooling them would produce an average describing no real patient.

The harm signal. Venous thromboembolism roughly doubled, on 48 events against 26, with an interval excluding 1. On a small number of events in a trial not primarily designed to measure thrombosis, this is a finding to take seriously without over-quantifying. Note the asymmetry from chapter 09: modified ITT dilutes benefit but also understates harm, since only patients who received the drug can be harmed by it.

LessonPopulation, dose and timing are not details — they are the intervention. A true result travels only as far as the population, regimen and time window in which it was established, and the commonest error in applying evidence is not statistical but a mismatch of one of those three. And a null primary outcome does not make a trial uninformative about harm: HALT-IT's most consequential number for practice is the one in its safety analysis, not its primary result.
05 · Derivation, validation, and reading a rule-out rule
Identification of children at very low risk of clinically-important brain injuries after head trauma: a prospective cohort study
Kuppermann N, Holmes JF, Dayan PS, Hoyle JD Jr, Atabaki SM, Holubkov R, Nadel FM, Monroe D, Stanley RM, Borgialli DA, Badawy MK, Schunk JE, Quayle KS, Mahajan P, Lichenstein R, Lillis KA, Tunik MG, Jacobs ES, Callahan JM, Gorelick MH, Glass TF, Lee LK, Bachman MC, Cooper A, Powell EC, Gerardi MJ, Melville KA, Muizelaar JP, Wisner DH, Zuspan SJ, Dean JM, Wootton-Gorges SL; Pediatric Emergency Care Applied Research Network (PECARN). Lancet 2009;374(9696):1160–70. PMID 19758692.
Design
Prospective cohort Derivation and validation of age-specific prediction rules, 25 North American emergency departments.
Population
42,412 children under 18 presenting within 24 h of head trauma with GCS 14–15. Derivation and validation sets: 8502 and 2216 aged under 2 years; 25,283 and 6411 aged 2 and over.
Outcome
Clinically important traumatic brain injury (ciTBI), defined as death from traumatic brain injury, neurosurgery, intubation for more than 24 h, or hospital admission for two or more nights.
Event rate
CT obtained in 14,969 (35.3 %). ciTBI occurred in 376 (0.9 %); 60 (0.1 %) underwent neurosurgery.
Under 2 y
Rule: normal mental status, no scalp haematoma except frontal, no loss of consciousness or LOC under 5 s, non-severe injury mechanism, no palpable skull fracture, acting normally per the parents. Validation NPV 1176/1176 = 100.0 % (95 % CI 99.7–100.0), sensitivity 25/25 = 100 % (86.3–100.0). 167 of 694 CT-imaged children (24.1 %) fell in this low-risk group.
2 y and over
Rule: normal mental status, no loss of consciousness, no vomiting, non-severe injury mechanism, no signs of basilar skull fracture, no severe headache. Validation NPV 3798/3800 = 99.95 % (99.81–99.99), sensitivity 61/63 = 96.8 % (89.0–99.6). 446 of 2223 CT-imaged children (20.1 %) fell in this low-risk group. Neither rule missed neurosurgery in the validation populations.

Why the outcome definition is the most important line in the paper. ciTBI is defined by what happens to the child, not by what the scan shows. That is a deliberate choice, and it is what makes the rule usable: a tool built to detect any abnormality on CT would inevitably recommend scanning for findings that change nothing, while accepting more radiation in a population where radiation risk is the entire motivation. When a decision rule seems to tolerate "missing" injuries, check its outcome definition before objecting — the misses may be, by design, injuries that needed nothing.

Read the eligibility criterion as a hard boundary. The cohort is GCS 14–15 within 24 h of injury. The rule says nothing whatever about a child with GCS 13, and applying it there is not a cautious extension but a use outside validation. This is the single commonest bedside misapplication of any decision rule.

Read the negative predictive value with its prevalence. ciTBI occurred in 0.9 % of the whole cohort, so a tool that ruled out everybody would already achieve an NPV above 99 %. The NPV is impressive because it is paired with high sensitivity in a population where the outcome is genuinely rare — and because the confidence interval is reported. Note the 2-and-over rule's sensitivity: 61 of 63, with an interval down to 89.0 %. Two children with ciTBI were in the low-risk group, neither needing neurosurgery. A rule-out tool with a stated, small, characterised miss rate is a usable tool; one presenting itself as infallible is not.

Derivation and validation in one paper. The rules were derived in one part of the cohort and applied, fixed, to patients not used in building them. That is real validation, not a resubstitution estimate — but it is internal validation within one research network in North America, so transportability to other systems and case mixes is a separate question answered by separate work, and the rule's calibration in a lower-prevalence setting is not established by this paper.

LessonAsk which stage of its life a prediction rule is in, and what its outcome actually was. Derivation-only performance is optimistic by construction; validation in new patients is the minimum for use; impact on patient outcomes and resource use is a further question again. And the boundaries of the validated population — here GCS 14–15, under 18, within 24 h — are the boundaries of the rule, no matter how reasonable the extrapolation feels at 3 a.m.
06 · The interval the p-value hides
Effect of a resuscitation strategy targeting peripheral perfusion status vs serum lactate levels on 28-day mortality among patients with septic shock: the ANDROMEDA-SHOCK randomized clinical trial
Hernández G, Ospina-Tascón GA, Damiani LP, Estenssoro E, Dubin A, Hurtado J, Friedman G, Castro R, Alegría L, Teboul JL, Cecconi M, Ferri G, Jibaja M, Pairumani R, Fernández P, Barahona D, Granda-Luna V, Cavalcanti AB, Bakker J; The ANDROMEDA SHOCK Investigators and the Latin America Intensive Care Network (LIVEN). JAMA 2019;321(7):654–664. PMID 30772908.
Design
RCT Multicentre randomised trial, 28 intensive care units in 5 countries, March 2017 to March 2018.
Population
424 patients with septic shock (mean age 63 years; 226 [53 %] women); 416 (98 %) completed the trial.
Intervention
Step-by-step resuscitation protocol targeting normalisation of capillary refill time (n = 212) versus normalising or decreasing lactate by more than 20 % per 2 h (n = 212), over an 8-hour intervention period.
Primary
All-cause mortality at 28 days.
Result
74 (34.9 %) vs 92 (43.4 %) died; hazard ratio 0.75 (95 % CI 0.55 to 1.02); p = 0.06; risk difference −8.5 % (95 % CI −18.2 % to 1.2 %).
Secondary
Less organ dysfunction at 72 h — mean SOFA 5.6 (SD 4.3) vs 6.6 (SD 4.7); mean difference −1.00 (95 % CI −1.97 to −0.02); p = 0.045. No significant differences in the other 6 secondary outcomes. No protocol-related serious adverse reactions confirmed.
Conclusion
A strategy targeting normalisation of capillary refill time did not reduce all-cause 28-day mortality. Registered NCT03078712.

What "p = 0.06" is being asked to carry here. The authors' conclusion — that the strategy did not reduce mortality — is the correct formal statement, because the trial did not meet its threshold. But read the interval alongside it. A hazard ratio of 0.75 with bounds of 0.55 and 1.02 is compatible with a 45 % relative reduction in death and with a 2 % relative increase. The absolute risk difference is −8.5 %, with an interval from −18.2 % to +1.2 %. Almost all of that interval lies in the region of substantial benefit. Summarising it as "capillary refill time made no difference" discards most of the information the trial produced, and the honest reading is: an inconclusive trial whose point estimate favours the intervention and which was not large enough to settle the question.

The threshold is a convention, not a finding. Had 3 more deaths occurred in the lactate arm, or 3 fewer in the perfusion arm, the result would have crossed p = 0.05 on a two-sided Fisher exact test of these counts, and the conclusion sentence in every subsequent citation would have been the opposite. Enter the counts — 74/212 against 92/212 — into the fragility calculator, which is preloaded with them, and see how few patients separate the two narratives. That fragility is the substance; the threshold is bookkeeping. (Fisher's exact test on the raw proportions gives p = 0.0906, not the 0.06 the trial reported from its time-to-event analysis — the two tests are asking slightly different questions of the same patients, which is itself worth noticing.)

And resist the opposite error. None of this makes the trial positive. A point estimate favouring the intervention on an inconclusive trial is a reason to want a larger trial, not a reason to change practice — and the SOFA difference of 1.00 point, with an upper bound of −0.02, sits on the edge of its own threshold and on a scale whose minimal clinically important difference is not established by this paper. Six other secondary outcomes did not differ. What the trial firmly establishes is that a bedside clinical target was not detectably worse than a laboratory one, which is itself a useful thing to know in a department without rapid lactate access.

LessonNever let a p-value be the summary of a trial. "Statistically significant" and "clinically important" are independent properties, and a result can be either, both or neither. Report the estimate, its interval, and what each end of the interval would mean for a patient — then, if you must, mention the p-value. A trial that fails to reject the null has not demonstrated equivalence; unless its interval is narrow and its bounds trivial, it has demonstrated that nobody yet knows.
Sources
How the figures on this tab were obtained

Every number, confidence interval, p-value, author list, journal, volume and page reference on this tab was taken from the paper's own record retrieved from PubMed via the NCBI E-utilities interface on 26 September 2026, and nothing has been added from memory or from secondary summaries. Where a figure a reader might expect is absent — a number needed to treat, a fragility index, an absolute risk difference not reported by the authors — it is absent because the source does not state it, and the calculators on this site will compute it from the counts that the source does state.

These appraisals are teaching material. They describe what each trial reported and what method lesson it illustrates; they are not recommendations, and none of the six should be used to guide a treatment decision without reading the paper itself and the current guidance for the condition.

  1. CRASH-2 trial collaborators; Shakur H, Roberts I, et al. Effects of tranexamic acid on death, vascular occlusive events, and blood transfusion in trauma patients with significant haemorrhage (CRASH-2): a randomised, placebo-controlled trial. Lancet 2010;376(9734):23–32. PMID 20554319.
  2. CRASH-2 collaborators; Roberts I, Shakur H, et al. The importance of early treatment with tranexamic acid in bleeding trauma patients: an exploratory analysis of the CRASH-2 randomised controlled trial. Lancet 2011;377(9771):1096–1101. PMID 21439633.
  3. Holcomb JB, Tilley BC, Baraniuk S, et al; PROPPR Study Group. Transfusion of plasma, platelets, and red blood cells in a 1:1:1 vs a 1:1:2 ratio and mortality in patients with severe trauma: the PROPPR randomized clinical trial. JAMA 2015;313(5):471–82. PMID 25647203.
  4. HALT-IT Trial Collaborators. Effects of a high-dose 24-h infusion of tranexamic acid on death and thromboembolic events in patients with acute gastrointestinal bleeding (HALT-IT): an international randomised, double-blind, placebo-controlled trial. Lancet 2020;395(10241):1927–1936. PMID 32563378.
  5. Kuppermann N, Holmes JF, Dayan PS, et al; PECARN. Identification of children at very low risk of clinically-important brain injuries after head trauma: a prospective cohort study. Lancet 2009;374(9696):1160–70. PMID 19758692.
  6. Hernández G, Ospina-Tascón GA, Damiani LP, et al. Effect of a resuscitation strategy targeting peripheral perfusion status vs serum lactate levels on 28-day mortality among patients with septic shock: the ANDROMEDA-SHOCK randomized clinical trial. JAMA 2019;321(7):654–664. PMID 30772908.

The eight pictures that carry most papers

A results section is usually two or three figures and some prose explaining them. Recognising the figure and knowing what it conceals is a large fraction of appraisal, and it is quicker to learn than any of the arithmetic.

Each panel below says what you are looking at, then what it hides. Every figure here is computed, not drawn — from a published trial's own numbers where one exists, or from simulated data with a stated model where it does not, and each is labelled which. A funnel plot sketched by hand to look asymmetric would be invented data, so none of these were.

01 · Trial flow real data
The CONSORT flow diagram
CONSORT flow diagram — the four numbers to findReal CRASH-2 figures · the losses here are what make its ITT credibleRandomised20,211 patientsAllocated tranexamic acid10,096 allocatedAllocated placebo10,115 allocatedAnalysed10,060 analysed36 excluded (0.4 %)Analysed10,067 analysed48 excluded (0.5 %)What to checkrandomised vsanalysed, in EACHarm separately.Are the lossessmall, and arethey balanced?

What it is. The map of every patient from screening to analysis, required by the CONSORT reporting guideline. It is the first figure in most trial reports and the one readers skip most reliably.

What to take from it. Two numbers per arm: how many were randomised, and how many were analysed. The difference between them, arm by arm, is the entire attrition story. CRASH-2 lost 36 of 10,096 and 48 of 10,115 — under half a per cent, and balanced — which is why its intention-to-treat analysis can be taken at face value.

What it hides. Losses that are balanced in number can still be biased in kind if the reasons differ between arms, and the diagram gives counts rather than reasons. Look for the reasons in the text. And a large gap between screened and randomised does not bias the result but does narrow it — that gap is the difference between the trial's population and yours.
02 · Describing one variable simulated data
Histogram, box plot and summary — the same 40 numbers three ways
One skewed dataset, three ways of looking at it40 simulated door-to-CT times in minutes · the same 40 numbers in every panelThe long tail drags the mean above the median. That gap IS the skew.2550100150200minutesmedian 38.5mean 56.5BOX PLOT — the same numbers on the same scaleQ1 29.8Q3 62.2median5 outliers (> Q3 + 1.5×IQR)Box = IQR · line = median · whiskers = last point within 1.5×IQRSUMMARYn40mean56.5 minmedian38.5 minSD44.7SEM7.06range18–214Q1–Q329.8–62.2IQR32.5Skewed, so report themedian and IQR — notthe mean and SD.

What it is. A histogram bins the values and shows the shape of the distribution. A box plot compresses that shape to five numbers: the median line, the box spanning the interquartile range, whiskers out to the last point within 1.5 × IQR, and individual dots for anything beyond.

What to take from it. These are skewed data — a long right tail, as almost all emergency-department time data have. The mean (56.5 min) sits well above the median (38.5 min), and that gap is the skew. Reporting "mean 56.5, SD 44.7" would describe a patient experience almost nobody had: mean minus two SD is negative.

What it hides. A box plot cannot show bimodality — two separate peaks can produce exactly the same five numbers as one broad hump, and the box plot will look unremarkable. Only the histogram, or a plot of the individual points, reveals it. If a paper gives you only box plots for a variable you would expect to be bimodal (two pathways, two operators, day versus night), that is a real limitation, not a quibble.

Paste your own series into the descriptive statistics calculator, which is preloaded with exactly these 40 numbers and draws both panels live.

03 · The reference distribution idealised
The normal distribution and the 68–95–99.7 rule
The normal distribution, and what one SD buys youIdealised curve · not data-3 SD-2 SD-1 SDmean+1 SD+2 SD+3 SD68 %95 %99.7 %mean = median = modeonly when it really is normalSkew moves the meanaway from the median

What it is. The symmetric bell curve that a great deal of statistical machinery assumes. It is drawn here from the equation, not from data.

What to take from it. The fixed proportions are the whole reason the standard deviation is a useful summary: about 68 % of observations lie within one SD of the mean, 95 % within two, 99.7 % within three. This is also where laboratory reference ranges come from — the central 95 % of a healthy population — and therefore why one healthy person in twenty falls outside the range for any given test.

What it hides. The proportions hold only if the distribution really is normal. On skewed data, "mean ± 2 SD" produces impossible values, which is the fastest way to catch a summary that should have been a median and IQR. Physiological variables are frequently not normal, and sample size does not fix that — a large sample of skewed data is a precisely measured skewed distribution.
04 · Comparing effects real data
The forest plot
CRASH-2: death due to bleeding, by time to treatmentCounts from the trial; intervals recomputed and matching those published.SubgroupTXA / placeboRR (95% CI)0.50.711.52risk ratio (log scale)← favours tranexamic acidfavours placebo →Overall489/10060 vs 574/100670.85 (0.76–0.96)≤ 1 h198/3747 vs 286/37040.68 (0.57–0.82)1–3 h147/3037 vs 184/29960.79 (0.64–0.97)> 3 h144/3272 vs 103/33621.44 (1.12–1.84)

What it is. One row per study or per subgroup: a marker at the point estimate, a horizontal line for its confidence interval, and a vertical line at no effect. Marker area is proportional to weight, so the eye is drawn to the studies carrying the result. A pooled estimate, where there is one, is drawn as a diamond whose width is its interval.

What to take from it. This one uses CRASH-2's own counts for death due to bleeding, split by time to treatment. The reading is immediate in a way the numbers alone are not: the two early bands sit left of the line, the late band sits clearly right of it, and the intervals of the first and last do not overlap. That visual separation is the interaction described in chapter 14.

What it hides. A forest plot makes subgroups look like separate studies, which they are not — randomisation guarantees comparability of the trial as a whole, not within a stratum defined after the fact. It also invites you to read significance off whether each line crosses the vertical, which is precisely the wrong way to assess a subgroup: the question is whether the rows differ from each other, and that needs a formal interaction test the plot cannot show. A log scale is conventional and necessary for ratios — check the axis, because on a linear scale a halving and a doubling look nothing alike.
05 · Publication bias simulated data
The funnel plot
Symmetricall 58 simulated studies plotted0.512odds ratio0.050.200.40standard errorlarger, more precise studiesAsymmetricsame data, 16 small null studies removed0.512odds ratio0.050.200.40standard errorlarger, more precise studies

What it is. Every study in a meta-analysis plotted as effect size against precision, with the most precise studies at the top. Because small studies scatter more, an unbiased set of studies forms a symmetric inverted funnel around the true effect.

What to take from it. Both panels are the same simulated studies; the right-hand one simply has some small studies that found nothing removed, exactly as failure to publish would remove them. The result is a visible notch in the bottom corner on the null side, and a pooled estimate that would be pulled away from the truth. Asymmetry is a prompt to ask where the missing studies went.

What it hides — and this is the commonly tested point. Funnel-plot asymmetry is not synonymous with publication bias. It also arises from genuine small-study effects (small trials are often conducted in higher-risk patients, or with more intensive protocols), from differences in methodological quality correlating with size, and from pure chance. Egger's and Begg's tests formalise the impression but are badly underpowered below about ten studies, so a symmetric-looking funnel with eight studies is no reassurance whatever. Treat asymmetry as a question, never a verdict.
06 · Time to event simulated data
Kaplan–Meier survival curves
Proportional hazardssimulated · curves never cross025507510006121824months% survivingAt risk220174139114100220148995940Crossing curvessimulated · log-rank goes half-blind here025507510006121824months% survivingAt risk220174139114100220157153144140

What it is. The proportion still event-free over time, as a step function — each step down is an event, and censored patients (lost to follow-up, or still event-free when the study ended) leave the denominator without causing a step.

What to take from it. Read the numbers at risk underneath before anything else. They are the reason the right-hand end of any survival plot deserves suspicion: that is where the fewest patients remain and the uncertainty is widest, and it is exactly where curves are most often described as "separating".

What it hides. The left panel has proportional hazards; the right panel's curves cross, and that changes what the statistics mean. A single hazard ratio averaged over the whole follow-up describes neither period, and the log-rank test loses power when curves cross because an early disadvantage and a late advantage cancel inside the statistic — two obviously different curves can return a comfortably non-significant p. See chapter 18. Kaplan–Meier also assumes censoring is non-informative; when patients are lost because they deteriorated, the curve is optimistic, and nothing on the plot reveals it.
07 · Diagnostic performance simulated data
The ROC curve and the Youden point
ROC curve and the Youden pointSimulated biomarker · 300 with disease, 700 without000.250.250.50.50.750.7511Youden: cut-off 9.3sens 81% · spec 77%1 − specificity (false positives)sensitivityAUC 0.86chanceThe two distributions the curve is built fromcut-offno diseasediseaseevery cut-off is one point on the curve

What it is. Sensitivity against 1 − specificity, with one point for every possible cut-off of a continuous test. The diagonal is chance; the top-left corner is perfection. The area under the curve is the probability that a random patient with the disease scores higher than a random patient without it.

What to take from it. The two overlapping distributions on the left are what the curve is actually summarising. Where they separate, the curve bulges towards the corner; where they overlap, it collapses onto the diagonal. Each point on the curve is a threshold, so the curve is a menu of trade-offs rather than a single property of the test.

What it hides. The Youden point marked here maximises sensitivity plus specificity, which silently assumes a false negative and a false positive cost the same. In emergency medicine they almost never do, so the statistically optimal cut-off is frequently the clinically wrong one. AUC is also threshold-free — a good AUC does not guarantee that any usable cut-off exists — and it says nothing about calibration, whether a predicted 5 % risk corresponds to an observed 5 %. Finally, a cut-off chosen in the same dataset that reports its performance is overfitted and will not reproduce.
08 · Comparing two methods simulated data
Scatter plot versus Bland–Altman
Correlation is not agreementSimulated: a new weight-estimation method against a reference, 90 children202040406060reference weight (kg)new method (kg)r = 0.990looks superbline of identitybiasupper limitlower limit204060+0+5+10mean of the two methods (kg)Bland–Altman: difference vs meanSame data, both panels. The new method reads about 7.1 kg high on average and the error grows with size:invisible in r = 0.990, unmissable on the right. Bias +7.1 kg, limits of agreement +1.8 to +12.4 kg.

What it is. On the left, a conventional scatter of a new method against a reference with a fitted line and a correlation coefficient. On the right, the same data plotted as Bland–Altman: the difference between the two methods against their mean, with the mean difference (bias) and the limits of agreement.

What to take from it. The correlation is 0.990, which looks conclusive and answers the wrong question. Correlation measures whether two things move together; it is unchanged if one method reads a fixed amount high, or a fixed percentage high. The Bland–Altman plot asks whether they agree, and shows that the new method reads about 7 kg high, with an error that grows as the child gets bigger — a fan-shaped spread that is completely invisible on the left.

What it hides. Limits of agreement are a description, not a verdict: somebody has to decide in advance how much disagreement would be clinically acceptable, and a paper that reports limits without that pre-stated threshold has left the judgement to you. The limits themselves have confidence intervals that are rarely drawn, and they are meaningless if the differences are not roughly normally distributed or if the spread changes with magnitude — as it does here, which is the case for a log transformation or for reporting percentage rather than absolute differences.
Provenance
Where each figure's numbers came from
FigureDataSource
CONSORT flowRealCRASH-2 randomisation and analysis counts, PMID 20554319
Forest plotRealCRASH-2 death due to bleeding overall (PMID 20554319) and by time band (PMID 21439633). Intervals recomputed from the counts and confirmed to reproduce those published
Histogram / box plotSimulated40 invented door-to-CT times, listed in full in the calculator
Normal curveIdealisedDrawn from the equation; not data
Funnel plotSimulated58 studies drawn around a true odds ratio of 0.80, with study size sampled continuously; the right panel removes small studies whose estimate fell on the null side
Kaplan–MeierSimulated220 per arm, exponential survival; the crossing panel adds an early high-hazard fraction
ROC curveSimulated300 with disease and 700 without, drawn from two overlapping normal distributions
Bland–AltmanSimulated90 paired measurements with a proportional bias deliberately built in

Every figure is generated by a script held with the page source rather than drawn by hand, so each one is reproducible and can be checked against the numbers above. The simulated figures exist because no real dataset can be shown that is guaranteed to illustrate one specific artefact cleanly; they are labelled as simulated on the figure itself so they can never be mistaken for evidence about a real treatment.

About this page

A critical appraisal and evidence-based medicine curriculum written for emergency and acute care clinicians — the twenty-three chapters, six calculators, five appraisal checklists, six worked papers and eight annotated figures that between them cover what a clinician needs in order to read a paper and decide whether to act on it.

It is part of ResusDoc, a free set of UK emergency medicine tools built and maintained by Dr Nirmalya Hore, an NHS emergency physician. There are no adverts and no tracking beyond aggregate site analytics; nothing you type into a calculator leaves your device.

This is an educational resource about method. It is not clinical guidance. The six trials on the Worked papers tab were chosen because each teaches an appraisal trap cleanly, not because they represent current practice, and several were chosen precisely because their clinical message is more complicated than their headline. Nothing here should be used to make a treatment decision.

The checklists are teaching summaries, not the instruments. The questions under each are phrased for a reader forming a judgement, and they compress the substance of the named tools. For a formal risk-of-bias assessment or a submission, obtain the current official version of RoB 2, ROBINS-I, QUADAS-2, PROBAST, AMSTAR 2, AGREE II, CASP, CONSORT, STROBE, PRISMA, STARD, TRIPOD or CHEERS from its own publisher. Reporting guidelines are revised, and a summary written once will drift.

The calculators implement standard published methods — Wilson intervals for proportions, the log method for risk ratios, Woolf's method for odds ratios, Wald intervals for risk differences, a two-sided Fisher exact test for the fragility index, the normal-approximation formula for two-proportion sample size, and for the descriptive calculator the n−1 sample variance, quartiles by linear interpolation between order statistics, and the adjusted Fisher–Pearson skewness coefficient. They are intended for appraisal and teaching, and are not a substitute for a statistical package in the analysis of real data. Each one names the method it used alongside its output so that a result can be checked.

The figures are computed, not drawn. Every figure on the Figures tab is generated by a script kept with the page source, from either a published trial's own verified numbers or a stated simulation with a fixed seed, and each is labelled on the figure itself as real, simulated or idealised. None was sketched to look a certain way, because a plot drawn to illustrate an artefact is invented data. The provenance of all eight is tabulated at the foot of that tab.

What is deliberately not here. The mathematics behind any of the tests; the computation of a posterior distribution, as opposed to how to read one; machine-learning model evaluation; and formal instruction in conducting research rather than reading it. Each is a separate subject, and a page that covered all of them would teach none of them.

Every figure on the Worked papers tab was taken from the paper's own PubMed record, retrieved through the NCBI E-utilities interface on 26 September 2026 — including each author list, journal, volume, page range and PMID. No figure, byline or citation on that tab comes from memory or from a secondary summary, and where the source does not state a number it is not stated here either.

The two figures on the Figures tab that use real data — the CONSORT flow diagram and the forest plot — use CRASH-2's own counts, and the forest plot's confidence intervals were recomputed from those counts and confirmed to reproduce the published ones before the figure was used. The other six are simulated or idealised and say so on their face.

The curriculum chapters describe generally accepted appraisal method as taught in clinical epidemiology, and name the reporting guidelines and risk-of-bias tools by their published names. Where a chapter illustrates a point with a real trial, the figures are the verified ones from the Worked papers tab and no others.

Review status: internally reviewed only. There has been no external review of this page's content, statistical or educational. If you find an error — in a formula, an interval, a quoted figure or a description of a method — it is worth reporting, and corrections are made rather than defended.

The page is built to be usable in a departmental teaching session without preparation.

  • Print collapses nothing — the print stylesheet expands every chapter, checklist and card and removes the navigation, so the whole thing prints as a handout.
  • Expand all in the header opens every section on the current tab, which is the fastest way to search the page with your browser's own find function.
  • Deep links work. Every tab and section has a hash, so #t14 opens the subgroups chapter, #c-fagan the pre-test to post-test calculator, #k-quadas the diagnostic checklist, #f-funnel the funnel plot and #pp-proppr the PROPPR appraisal. Those links can go straight into a teaching email or a rota message.
  • The calculators are preloaded with real published counts — CRASH-2 mortality in the effect calculator, ANDROMEDA-SHOCK mortality in the fragility calculator — so a session can begin by reproducing a published interval and then changing one number to see what it takes to change the conclusion.
  • The Figures tab is the fastest thing here to teach from: each panel says what the plot shows and then what it hides, and the two using real trial data are labelled as such.
  • Chapter 23 contains a ten-minute appraisal routine and a set of suggestions for running a journal club that ends in a recorded decision rather than a general air of scepticism.