The question a clinician actually needs to answer about a paper is narrow: would acting on this change what I do to the next patient through my door, and how badly could it be wrong? Almost everything below exists to answer that. The statistics are a means to it, not the point of it.
Three habits do most of the work. Read the methods before the results — the design fixes the ceiling on what the numbers can mean, and by the time you reach the abstract's conclusion you are being told what to think. Convert every relative number into an absolute one, because a 30 % risk reduction is a different intervention at a 20 % baseline risk than at a 0.2 % one. And ask what the trial's population was, because population mismatch, not statistics, is the commonest reason a true result does not apply to your patient.
The twenty-three chapters here run from describing a single variable to GRADE. The Calculators tab does the arithmetic, the Figures tab shows what the plots in a results section are hiding, the Worked papers tab appraises six real trials, each chosen because it teaches one trap cleanly.
A paper read without a question in mind is a paper that will tell you its own conclusion. Writing the question down first is not a formality — it is what makes the relevance judgement happen before you have been persuaded.
| Element | Asks | Where appraisal goes wrong |
|---|---|---|
| Population | Which patients? | The trial's population is narrower than the specialty that adopts it. Severity, age, setting and time-from-onset are the usual mismatches. |
| Intervention | Exactly what, at what dose, how soon? | A regimen gets simplified in the retelling. A 1 g bolus plus an 8-hour infusion becomes "give tranexamic acid". |
| Comparator | Against what? | Placebo answers "does it work"; usual care answers "should I switch". A trial against a straw-man comparator flatters the drug. |
| Outcome | Which outcome, measured when? | The outcome that made the headline is often not the primary outcome, and often not one a patient would notice. |
Then predict the answer. Before you read the results, commit to what you expect and how big an effect would change your practice. This single move protects against the two commonest failures at once: being impressed by a result you would have dismissed had it gone the other way, and calling a trial "negative" when it merely failed to find the implausibly large effect it was powered for.
The first fork is interventional versus observational: did the investigators allocate the exposure, or only watch it? Everything about what a study can claim follows from that. The second fork, within observational work, is the direction of sampling: did you sample on exposure and wait for outcomes (cohort), or sample on outcome and look back for exposure (case–control)?
| Design | Sampling | Answers well | Structural weakness |
|---|---|---|---|
| Randomised controlled trial | Allocated by chance | Does this intervention cause this outcome | Cost, narrow eligibility, often short follow-up, rarely powered for rare harms |
| Cluster RCT | Groups randomised | Systems, pathways, training | Effective sample size sits between the number of clusters and the number of patients, and is nearer the former the larger the clusters and the more alike patients within them; needs ICC adjustment in both the sample size and the analysis |
| Stepped-wedge | Sequential rollout | Roll-outs that cannot be withheld | Confounded by secular trend; needs time in the model |
| Prospective cohort | On exposure | Incidence, prognosis, multiple outcomes | Confounding; loss to follow-up |
| Retrospective cohort | On exposure, from records | Rare exposures, long latency | Data collected for another purpose; exposure misclassification |
| Case–control | On outcome | Rare outcomes, many exposures | Recall and selection bias; gives odds ratios, not risks |
| Cross-sectional | One time point | Prevalence, diagnostic accuracy | No temporality — cannot separate cause from consequence |
| Case series / report | Whoever appeared | Novelty, harm signals, the unexpected | No denominator, no comparison — cannot estimate effect at all |
| Diagnostic accuracy | Consecutive suspected cases | How well a test classifies | Reference-standard quality; spectrum and selection effects |
| Systematic review | Studies, not patients | Synthesis, precision, consistency | Inherits its constituents' bias; heterogeneity may be real |
The pyramid with RCTs near the top is a statement about average susceptibility to bias, not a ranking of individual papers. A small, unblinded, selectively reported RCT stopped early for benefit is weaker evidence than a large, carefully adjusted prospective cohort with complete follow-up. Judge the study, not its label.
An explanatory trial asks whether a thing can work under ideal conditions: tight eligibility, protocol-driven care, expert centres. A pragmatic trial asks whether it does work in the real world: broad eligibility, usual care as comparator, routinely collected outcomes. Pragmatic trials generalise better and tend to show smaller effects — which is usually the truth rather than a failure. When a pragmatic trial disappoints after an explanatory one impressed, the effect has not disappeared; it has been measured under conditions you actually work in.
Before any test is chosen or any p-value computed, a paper has to say what its data are. Almost every later decision — which summary to report, which test is legitimate, whether a mean is even meaningful — follows from the answer.
| Type | Example | Legitimate summary | Watch for |
|---|---|---|---|
| Nominal (categorical, unordered) | Blood group, presenting complaint, discharge destination | Counts and proportions | There is no average. A "mean blood group" is not a quantity. |
| Binary | Died / survived, admitted / discharged | Proportion, risk, odds | A special case of nominal, and the one most of this site is about. |
| Ordinal (ordered, unequal gaps) | Triage category, pain score 0–10, mRS, GCS | Median and IQR; proportions above a threshold | The gaps are not equal — the step from mRS 4 to 5 is not the step from 0 to 1. Means of ordinal scales are reported constantly and are hard to interpret. |
| Discrete count | Attendances per shift, number of seizures | Median and IQR, or a rate | Often skewed and bounded at zero. |
| Continuous | Age, lactate, door-to-CT time, sodium | Mean and SD if symmetric; median and IQR if not | Measurement precision, and whether it was later chopped into categories. |
In a perfectly symmetric distribution the three coincide. The gap between mean and median is itself a measure of skew — if a paper reports both and they differ substantially, the data are skewed whatever else the text says.
| Measure | What it is | Use when |
|---|---|---|
| Range | Lowest to highest | Almost never on its own — it depends entirely on the two most extreme patients and grows with sample size |
| Interquartile range (IQR) | 25th to 75th centile — the middle half of the data | Skewed data, ordinal data, anything with outliers. Report it as two numbers (Q1–Q3), not as their difference alone |
| Variance | Mean squared deviation from the mean | Rarely reported directly — its units are squared. The worked dataset on the Figures tab has an SD of 44.7 minutes and therefore a variance of 1,995 minutes², a number that means nothing to a reader on sight |
| Standard deviation (SD) | Square root of the variance, so back in the original units | Symmetric, roughly normal data |
A standard deviation is only interpretable if the distribution is roughly symmetric, because its usefulness comes entirely from the normal distribution's fixed proportions: about 68 % of observations within one SD of the mean, 95 % within two, and 99.7 % within three (strictly, 95 % falls within 1.96 SD — which is where the 1.96 in every confidence interval comes from). Quote an SD on a heavily skewed variable and "mean ± 2 SD" will include impossible values — negative door-to-CT times, negative lactates. That is the quickest way to spot a summary that should have been a median.
See the normal distribution figure and the skewed-dataset figure, where the same 40 numbers are shown as a histogram, a box plot and a summary table.
They answer different questions and are not interchangeable.
A centile is the value below which that percentage of observations fall; the median is the 50th. Paediatric growth and observation charts are built from them, which is why "below the 2nd centile" is a statement about a population, not a diagnosis.
Most laboratory reference ranges are the central 95 % of a healthy reference population — which has a consequence people rarely state out loud: 1 healthy person in 20 will fall outside the range for any given test, by construction. Order a panel of 20 independent tests on a well person and the expected number of "abnormal" results is one. A reference range is not a threshold for disease, and an isolated marginal abnormality on a broad panel is usually the arithmetic working as designed.
You will rarely choose a test yourself. You will constantly need to judge whether the one a paper used was reasonable — and that judgement reduces to three questions asked in order.
| Outcome | Two independent groups | Two paired measurements | Three or more groups |
|---|---|---|---|
| Continuous, roughly normal | Unpaired (Student's) t-test | Paired t-test | ANOVA |
| Continuous, skewed; or ordinal | Mann–Whitney U (Wilcoxon rank-sum) | Wilcoxon signed-rank | Kruskal–Wallis |
| Categorical / binary | Chi-square; Fisher's exact when expected counts are small | McNemar's test | Chi-square across the table |
| Time to event | Log-rank; Cox regression for adjustment | — | Log-rank across strata |
Two practical notes on that table. Fisher's exact is the safe choice whenever any expected cell count is small (the usual rule of thumb is under 5) — it is what the fragility calculator on this site uses. And ANOVA answers only "is there a difference somewhere"; it does not say which groups differ, which requires post-hoc comparisons with an explicit correction for multiplicity.
"Parametric" tests assume the data follow a particular distribution — usually the normal — and additionally, for the t-test and ANOVA, that group variances are similar and observations are independent. Non-parametric tests replace the values with their ranks and so assume much less.
| Question | Tool | Gives |
|---|---|---|
| Do these two move together? | Pearson's r (linear, continuous, normal); Spearman's rho (ranks — monotonic, ordinal, robust to outliers) | −1 to +1. r² is the proportion of variance shared |
| Can I predict one from the other, and by how much? | Linear regression | A slope in real units, with a confidence interval |
| Can this method replace that one? | Bland–Altman | Bias and limits of agreement, in real units |
Every test above returns a p-value, and a p-value is a statement about compatibility with a null hypothesis and nothing else. It does not give the size of the effect, its direction in clinically useful units, or its precision. A paper whose results section is a list of p-values has reported the least informative summary available. What you want, and should look for instead, is the estimate with its confidence interval — a difference in means, a difference in medians, a risk difference — in the units a patient would recognise. See chapter 11.
Bias is systematic error — it moves the estimate in a direction and does not shrink when you enrol more patients. That last clause is the whole reason bias matters more than imprecision: a bigger study fixes imprecision and entrenches bias. The five domains below follow a trial through its own timeline, which is also the order in which to look for them.
| Domain | When | What happens | Look for |
|---|---|---|---|
| Selection bias | At allocation | The groups differ at baseline in ways that affect the outcome | How the sequence was generated; whether allocation was concealed; the baseline table |
| Performance bias | During the trial | The groups get different co-interventions or attention | Blinding of patients and clinicians; protocol for concomitant care; crossover |
| Detection bias | At measurement | Outcomes are sought or scored differently between groups | Blinded outcome assessment; adjudication committee; objectivity of the outcome |
| Attrition bias | At follow-up | Losses differ between groups, or from those who stay | The CONSORT flow diagram; numbers analysed versus randomised; how missing data were handled |
| Reporting bias | At publication | Outcomes or whole studies are selectively reported | The registered protocol; whether the primary outcome changed; whether all pre-specified outcomes appear |
A confounder is associated with the exposure, is a cause of the outcome independently of the exposure, and is not simply a step on the causal path between them. That third clause matters: adjusting for a mediator removes part of the effect you are trying to measure, and adjusting for a collider — a common consequence of exposure and outcome — can manufacture an association out of nothing.
| Tool | Handles | Residual problem |
|---|---|---|
| Randomisation | All confounders, known and unknown | Only in expectation — small trials can still be unbalanced by chance |
| Restriction / matching | The variable you chose | Loses generalisability; over-matching can remove real effect |
| Stratification | A few categorical variables | Runs out of patients quickly |
| Multivariable regression | Several measured confounders | Needs the right variables, correctly measured, with a sensible model form |
| Propensity score | Many measured confounders at once | Same dependence on measured variables — it is bookkeeping, not magic |
| Instrumental variable / Mendelian randomisation | Unmeasured confounding, in principle | Requires an instrument that is genuinely unrelated to the outcome except through exposure |
Table 2 fallacy. The other coefficients in a multivariable model are not each a valid causal estimate for their own variable; the model was specified to estimate one exposure effect, and the covariates' coefficients carry a mixture of direct and confounded signal. A paper that presents its whole adjustment table as a list of independent risk factors is over-reading its own model.
These three are routinely conflated, and they protect against different failures at different moments.
Baseline tables and their misuse. The baseline table exists to let you judge whether randomisation produced comparable groups, and to see whom the trial enrolled. It should not carry p-values: in a properly randomised trial any baseline difference is by definition chance, so testing it answers a question nobody asked. What you want instead is the size of the imbalance and whether the imbalanced variable is prognostically important.
Outcome choice is where a trial decides how useful it is allowed to be, and it is decided before a single patient is enrolled.
A surrogate stands in for what you care about: lactate clearance for survival, radiographic union for function, viral load for illness. Surrogates make trials smaller and faster, and they mislead whenever the intervention affects the marker by a route that does not run through the outcome. The question to ask is not "is this surrogate correlated with the outcome" — it usually is — but "do interventions that move this surrogate reliably move the outcome?" That is a far higher bar, and most surrogates fail it.
Composites (death, MI or urgent revascularisation; death or dependency) buy statistical power by collecting more events. They are legitimate when the components are of similar importance to the patient, of similar frequency, and plausibly moved in the same direction by the intervention. They mislead when a frequent, mild, discretionary component — usually something like "unplanned re-attendance" — drives the whole result while the component anyone cares about does not move. Always look for the component-by-component breakdown; if it is absent, treat the composite as unresolved.
Mortality, function, symptom burden, and days alive and out of hospital are outcomes patients recognise. Patient-reported outcome measures (PROMs) capture them directly, at the cost of needing a validated instrument, a minimal clinically important difference (MCID) established independently of the trial, and an honest account of missing questionnaires. A statistically significant change on a scale is meaningless without knowing the MCID, and quoting one from the same trial that found the effect is circular.
| Population | Who is analysed | Preserves randomisation | Answers |
|---|---|---|---|
| Intention-to-treat (ITT) | Everyone randomised, in their allocated group, whatever happened next | Yes | What happens if I adopt this policy |
| Modified ITT (mITT) | ITT minus a pre-defined group, e.g. those who never received a dose | Partly — depends entirely on the exclusion rule | Somewhere between the two; needs scrutiny |
| Per-protocol | Only those who completed treatment as specified | No | What happens in patients who tolerate and comply — a self-selected group |
| As-treated | Grouped by what they actually received | No | An observational question inside a trial |
Why ITT is the default. Dropping non-compliers destroys the one thing randomisation gave you, because non-compliance is not random: people stop drugs because of side effects, because they deteriorate, or because they recover. Per-protocol analysis therefore compares a filtered intervention group with an unfiltered control group, and the filter is prognostic. ITT preserves the comparison and answers the question a clinician actually faces, which is "if I prescribe this, what happens" — including the patients who will not take it.
Reading the numbers-analysed line is not pedantry. CRASH-2 randomised 20,211 patients and analysed 10,060 of 10,096 and 10,067 of 10,115 — losses of well under 1 % in both arms, which is why its ITT analysis is credible at face value. HALT-IT states the opposite arrangement plainly: its primary analysis excluded patients who received neither dose and those without outcome data, which is a modified ITT, and the paper says so in its methods rather than leaving you to discover it. A trial that randomises 500 and analyses 380 has a problem regardless of what it calls the analysis.
Take a trial with a control event rate CER and an experimental event rate EER.
The relative measures answer "how much does this shift a person's risk", and they travel reasonably well between populations of different baseline risk. The absolute measures answer "how many people benefit", and they do not travel at all — they depend entirely on baseline risk. This is why a relative risk reduction quoted without a baseline risk is uninterpretable, and why it is the favourite number of anyone selling something.
| Setting | CER | EER | RRR | ARR | NNT |
|---|---|---|---|---|---|
| High-risk population | 20 % | 15 % | 25 % | 5 % | 20 |
| Low-risk population | 0.4 % | 0.3 % | 25 % | 0.1 % | 1000 |
Identical relative effect; fifty-fold difference in how many people you must treat. If the intervention carries a 1-in-200 risk of serious harm, it is clearly worth doing in the first row and clearly not in the second — and no statistical test will tell you that, because it is a judgement about magnitude, not significance.
An odds ratio is (a/b)/(c/d) on the 2×2 table. It is the natural output of logistic regression and of case–control studies, where risks cannot be estimated. Its trap is simple: the OR exaggerates the RR, and increasingly so as the outcome becomes more common. For an outcome affecting 2 % of controls, OR and RR are nearly identical; for one affecting 40 %, an OR of 2.0 corresponds to a risk ratio nearer 1.4. Treating a published OR as a risk ratio is one of the most common misreadings in clinical conversation.
A hazard ratio compares the instantaneous rate of the event between groups, using time-to-event data, and assumes proportional hazards — that the ratio is roughly constant over follow-up. When survival curves cross or diverge late, that assumption fails and the single HR hides the shape of what happened. Look at the curves and the numbers-at-risk row beneath them; late-separating curves with thin numbers at risk are where over-interpretation lives.
The treatment-effect calculator computes all of these with confidence intervals from raw event counts.
A p-value is the probability of observing data at least as extreme as yours if the null hypothesis were true and the model were correct. That is all it is. It is a statement about data given a hypothesis, not about a hypothesis given data, and the difference is not academic.
| p does not mean | Why not |
|---|---|
| The probability the null is true | That is an inverse probability and needs a prior. p cannot supply one. |
| The probability the result is a fluke | Same error, more casually phrased. |
| The size or importance of the effect | p mixes effect size with sample size. A trivial effect in a huge trial gives a tiny p. |
| p > 0.05 means no effect | Absence of evidence. A wide interval around a null point estimate is ignorance, not equivalence. |
| p = 0.049 and p = 0.051 differ meaningfully | The threshold is a convention, not a property of nature. |
| Two trials with p < 0.05 and p > 0.05 disagree | Their intervals usually overlap heavily. Compare estimates and intervals, not verdicts. |
A 95 % confidence interval is the range of effects reasonably compatible with the data. It carries everything the p-value carries (if it excludes the null, p < 0.05) and adds the two things a clinician needs: the plausible size of the effect, and the precision of the estimate. Read both of its ends and ask what each would mean clinically:
Test twenty independent true nulls at α = 0.05 and you expect one "significant" result. Trials handle this by nominating a single primary outcome, pre-specifying a small number of comparisons, and either splitting alpha (Bonferroni, hierarchical testing) or accepting that everything else is exploratory. PROPPR illustrates the honest handling: 23 pre-specified complications were examined and none differed, and the paper reports that as a safety statement rather than mining it for a finding.
The previous chapter ended on an awkward admission: the confidence interval does not mean what almost everyone takes it to mean, and the p-value answers a question nobody asked. Bayesian inference is the framework that answers the question people actually have — given these data, how likely is it that this treatment works? — at the cost of requiring you to state what you believed beforehand.
| 95 % confidence interval | 95 % credible interval | |
|---|---|---|
| Means | A procedure that captures the true value 95 % of the time across repeated experiments | Given the data and the prior, a 95 % probability that the true value lies inside |
| Requires | No prior | A stated prior |
| Can say | Nothing directly about the probability of benefit | A statement of the form "there is a 94 % probability that this treatment reduces mortality" — a real worked example follows below |
That last row is the practical difference, and it is the reason Bayesian re-analyses of intensive care and emergency trials have become common. A frequentist trial reporting p = 0.06 is obliged to say it failed to reject the null; a Bayesian analysis of the same patients can put a probability on benefit directly.
The honest objection to Bayesian analysis is that the prior can be chosen to get the answer you want. The honest answer is that a good Bayesian analysis makes the prior explicit and then shows how much it mattered. Look for:
Bayesian analysis is most valuable exactly where the frequentist framework is least informative: trials that miss the threshold, trials that stop early (where the framework has no multiplicity penalty, because the posterior does not care how many times you looked), small trials in rare conditions where a well-justified prior adds real information, and adaptive designs. It is not a way of rescuing a negative trial — applied to a result with a tight interval around no effect, it returns a posterior tightly centred on no effect, which is the same conclusion in different words.
Both frameworks describe the same data. They differ in what they are willing to put a probability on, and the practical instruction is the same for both: read the interval and the effect size, not the verdict.
Power is the probability of detecting an effect of a specified size if it is really there. It rises with the number of events (not patients), with the size of the effect sought, and with tolerance for a larger α. Four quantities are locked together — α, power, effect size and sample size — so fixing any three fixes the fourth.
Calculating power after the fact from the observed effect adds no information — it is a deterministic function of the p-value. If a trial was inconclusive, the confidence interval already says so, and says it better. Report the interval; do not compute observed power.
Repeatedly testing accumulating data inflates the type I error, so trials use alpha-spending (O'Brien–Fleming, Pocock) with a data monitoring committee and pre-specified stopping rules for benefit, harm or futility. Two practical consequences:
The fragility index is the number of patients whose outcome would have to change, in one arm, to move a significant result across p = 0.05. It is an intuition pump rather than a formal inference: a trial whose headline depends on two events is not the same evidence as one whose headline depends on eighty, even when both report p = 0.04. Its limitations are real — it depends on the test used, ignores effect size and time-to-event data, and can be gamed — so use it to calibrate confidence, not to overturn a result. The fragility calculator computes it with Fisher's exact test.
Subgroup analysis is where most over-interpretation in clinical medicine originates, and also where some of its most important findings came from. The difference between the two is knowable, and it is not whether the subgroup result is "significant".
Splitting a trial into subgroups and looking for which ones have p < 0.05 is close to guaranteed to mislead. Each subgroup is smaller and therefore less powered, so significance within subgroups tracks subgroup size as much as effect. The famous demonstration is the ISIS-2 astrological subgroup: aspirin appeared not to work in patients born under Gemini and Libra. Nobody believes that, and the statistics that produced it are the same statistics that produce the subgroup claims people do believe.
The question is not "is the effect significant in this subgroup" but "is the effect different between subgroups". That is a formal test for interaction (or, for an ordered variable, for trend), and it should be pre-specified, few in number, biologically motivated, and consistent across related outcomes.
One or two yeses is a hypothesis. Most of them, as with CRASH-2 timing, is something you can act on while still calling it what it is.
A superiority trial tries to show a difference and treats failure as failure. A non-inferiority trial tries to show the new treatment is not unacceptably worse than the old — usually because it is cheaper, safer, oral, or available at 3 a.m. Its logic is inverted, and so are several of its incentives.
Everything hangs on the pre-specified non-inferiority margin Δ: the largest loss of efficacy you are willing to accept. Set it narrow and the trial is honest and expensive. Set it wide and almost anything passes. The margin should be justified from the size of the comparator's own established effect over placebo — preserving a stated fraction of it — and not from what makes the sample size affordable. A non-inferiority trial with no justification for its margin cannot be appraised, only believed or not.
| Confidence interval versus Δ and the null | Conclusion |
|---|---|
| Entirely on the good side of Δ, crosses the null | Non-inferior, not superior |
| Entirely on the good side of Δ and excludes the null favourably | Non-inferior and superior |
| Crosses Δ | Non-inferiority not established — inconclusive, not "equivalent" |
| Entirely beyond Δ | Inferior |
Equivalence trials are two-sided non-inferiority: the interval must sit inside ±Δ. They are rare in clinical medicine outside bioequivalence. What is common, and wrong, is calling a failed superiority trial "equivalent" — you cannot switch design after the fact, because the failed trial never specified a margin.
| Disease present | Disease absent | |
|---|---|---|
| Test positive | TP (a) | FP (b) |
| Test negative | FN (c) | TN (d) |
Sensitivity and specificity are read down the columns — they condition on disease status, and are therefore roughly stable across populations of different prevalence (roughly, not exactly — see spectrum bias below). PPV and NPV are read across the rows, so they depend on prevalence and are not properties of the test at all. A d-dimer's NPV in a low-risk outpatient population and in a critical care unit are different numbers for the same assay. Quoting a PPV without the prevalence it was measured at is meaningless.
Likelihood ratios convert a pre-test probability into a post-test probability, which is what a clinician is actually doing when ordering a test. They work multiplicatively on odds:
| LR+ | Effect on probability | LR− | Effect |
|---|---|---|---|
| > 10 | Large increase — often conclusive | < 0.1 | Large decrease — often conclusive |
| 5–10 | Moderate increase | 0.1–0.2 | Moderate decrease |
| 2–5 | Small increase | 0.2–0.5 | Small decrease |
| 1–2 | Negligible | 0.5–1 | Negligible |
Everything above assumed the test gives a yes or a no. Most do not — troponin, lactate, d-dimer and every clinical score produce a number, and somebody has to choose where to cut it.
Take every possible cut-off in turn. At each one, compute sensitivity and 1 − specificity, and plot that pair as a point. Join the points and you have the receiver operating characteristic curve. Each point on it is a cut-off, so the curve is not a property the test has in the abstract — it is the complete menu of trade-offs the test offers.
The ROC figure shows a curve alongside the two distributions it was built from, which is the thing worth internalising: the curve is a summary of how much those two distributions overlap. Where they separate cleanly the curve bulges towards the corner; where they overlap heavily it collapses onto the diagonal.
The area under the curve (AUC, or c-statistic) summarises discrimination across all thresholds at once, and has a clean interpretation: the probability that a randomly chosen patient with the disease scores higher than a randomly chosen patient without it. 0.5 is chance; 1.0 is perfect. The conventional, and entirely arbitrary, descriptive bands are ~0.7–0.8 acceptable, 0.8–0.9 good, above 0.9 excellent.
The Youden index is the most commonly used rule: maximise (sensitivity + specificity − 1), which geometrically is the point on the curve furthest above the diagonal. It is marked on the figure. Its hidden assumption is the one that matters: Youden weights a false negative and a false positive equally.
In emergency medicine they are almost never equal. A missed subarachnoid haemorrhage and an unnecessary CT head are not comparable harms, and no arithmetic derived only from the curve knows that. So:
QUADAS-2 in the Checklists tab is the structured way to walk through these.
Emergency medicine runs on prediction rules, and almost every argument about one comes down to confusing the three stages of its life.
| Stage | What is done | What it establishes | What it does not |
|---|---|---|---|
| 1. Derivation | Candidate predictors collected; the rule is fitted to this dataset | That a combination of variables separates outcomes in these data | Anything about other patients. Performance here is optimistic by construction. |
| 2. Validation | The fixed rule is applied to new patients — internal (split/bootstrap), temporal, or external (different sites, different country) | That performance holds outside the derivation set | That anyone will use it, or that using it helps |
| 3. Impact | The rule is implemented and patient-level outcomes and resource use are measured, ideally in a cluster RCT | That deploying it changes practice for the better | — |
A prognostic study asks what happens to this patient over time, not whether an intervention helps. Its central technical problem is censoring: most patients have not had the event when the study ends, and simply excluding them would throw away nearly all the data.
For appraising a prognostic model specifically, the relevant reporting standard is TRIPOD, and the risk-of-bias tool is PROBAST; for a prognostic factor study, QUIPS.
A systematic review is a study whose unit of analysis is the study: a pre-registered question, an explicit search, duplicate screening and extraction, formal risk-of-bias assessment, and a stated synthesis plan. A meta-analysis is the optional statistical pooling step. A review can be excellent without pooling, and pooling without the rest is not a systematic review.
A fixed-effect model assumes one true effect and that differences between studies are sampling noise. A random-effects model assumes the true effect varies between studies and estimates the distribution's mean. Random effects gives wider intervals and weights small studies relatively more — which matters, because small studies are the ones most susceptible to bias. Neither model repairs heterogeneity; they make different assumptions about it.
| Statistic | Means | Caution |
|---|---|---|
| I² | Proportion of variability due to between-study differences rather than chance | Not an absolute measure; unreliable with few studies; a low I² with wide intervals proves nothing |
| τ² | Absolute between-study variance, on the effect scale | More informative than I² but rarely quoted |
| Cochran's Q, p | Test of homogeneity | Badly underpowered with few studies — a non-significant Q is not evidence of homogeneity |
| Prediction interval | Where the effect of a future study would plausibly fall | The most honest single summary of heterogeneity, and the least often reported |
Small studies with unremarkable results are less likely to be published, so meta-analyses over-represent positive small studies and overestimate effects. A funnel plot shows effect against precision; asymmetry suggests missing small negative studies. Egger's and Begg's tests formalise it. All of them are underpowered below roughly ten studies, and asymmetry has innocent explanations — genuine small-study effects, differing quality, or a real relationship between dose and population risk. Treat funnel asymmetry as a prompt, not a verdict.
GRADE separates two judgements that are constantly run together: how certain are we about the effect, and what should we therefore do. They can come apart in both directions, and understanding that is most of what a clinician needs from GRADE.
Certainty is rated per outcome, not per study and not per guideline. RCT bodies start high, observational bodies start low, and the rating then moves.
| Rate down for | Rate up for |
|---|---|
| Risk of bias in the contributing studies | Large effect size |
| Inconsistency (unexplained heterogeneity) | Dose–response gradient |
| Indirectness (population, intervention, comparator or outcome) | All plausible residual confounding would reduce the observed effect |
| Imprecision (intervals crossing decision thresholds, few events) | — |
| Publication bias | — |
The four resulting levels — high, moderate, low, very low — are statements about how likely further research is to change the estimate, not about study design alone.
A strong recommendation ("we recommend") means almost all informed patients would choose this, and it can reasonably become a quality standard. A conditional or weak recommendation ("we suggest") means the right choice depends on values and circumstances, and shared decision-making is the point rather than a courtesy. Strength depends on certainty plus the balance of benefits and harms, the variability of patient values, and resource use.
The tool is AGREE II, but three questions do most of the work: who wrote it and who paid (conflicts declared and managed, chair independence); is the link from evidence to recommendation traceable (does each recommendation carry its evidence and grade, or does the reader have to take it on trust); and is it current (a named review date, with a statement of what has changed). A guideline recommendation quoted without its number and its grade cannot be checked, which is why both belong in any citation of one.
Everything above assumes the paper is an honest report of what was done. That assumption is usually safe and is not free.
Prospective registration exists so that the primary outcome cannot be chosen after the results are known. Comparing the registry entry with the publication is a five-minute check that catches the single most consequential form of reporting bias: a primary outcome demoted, a secondary promoted, or a new outcome appearing. Registration numbers are given precisely so this is possible — CRASH-2 lists ISRCTN86750102 and NCT00375258; HALT-IT lists ISRCTN11225767 and NCT01658124; PROPPR lists NCT01545232; ANDROMEDA-SHOCK lists NCT03078712.
Spin is presentation that implies more than the data support without stating anything false. Its recurring forms:
A declared commercial interest does not invalidate a study, and its absence does not sanctify one — academic, intellectual and career conflicts are real and rarely declared. What matters is whether the funder had a structural role in the parts that can be steered: design, conduct, analysis, and the right to publish a negative result. Note that CRASH-2's funding included both public (UK NIHR Health Technology Assessment programme) and commercial (Pfizer) sources, alongside charitable funders, and that it reported a modest effect on a hard outcome — the combination is worth noting rather than either dismissing or ignoring.
Early positive results are, on average, exaggerated — a consequence of publication bias, small samples, flexible analysis and the winner's curse of reporting whichever estimate crossed the threshold. The practical implication is a prior, not a cynicism: expect a first striking result to shrink on replication, and be slower to adopt on a single trial than the trial's own confidence interval suggests. In a field where a single-centre trial can change practice within a month, that prior is the most useful thing on this page.
Paper mills, fabricated datasets and undisclosed generative-AI text are now part of the landscape, most heavily in low-barrier journals. Signals worth noticing: implausibly clean baseline tables, distributions too uniform, an author list with no traceable institutional link to the work, and reference lists containing citations that do not resolve. Check that a citation exists and says what it is claimed to say — the most reliably detectable defect in an unsound paper is not its statistics but its references.
Qualitative work answers questions quantitative work cannot: why a pathway is not followed, what a diagnosis means to a patient, how a team actually makes a decision under pressure. Appraising it against RCT criteria is a category error — sample size, blinding and generalisability are the wrong questions. The right ones are:
The reporting standards are COREQ (interviews and focus groups) and SRQR; CASP publishes a qualitative checklist, and GRADE-CERQual rates confidence in findings from qualitative evidence synthesis.
The question to ask is why the two strands were combined and how they were integrated. A paper that reports a survey and an interview study side by side, with no integration, is two papers sharing a title. Genuine mixed-methods work states its sequence (explanatory, exploratory or convergent) and shows where the strands informed each other.
An economic evaluation compares costs and consequences of alternatives. Key features to check: the perspective (health service versus societal — it changes which costs count), the time horizon and discount rate, whether the outcome is a cost-effectiveness ratio (cost per event avoided) or a cost–utility ratio (cost per QALY), the incremental cost-effectiveness ratio rather than average costs, and the sensitivity analysis — a probabilistic one with a cost-effectiveness acceptability curve, since a point ICER without uncertainty is not a result. The reporting standard is CHEERS. Be alert to models whose conclusion is driven by one poorly evidenced input; the sensitivity analysis is where that shows, which is why it is the section to read first.
This is the sequence to use on a paper someone has just sent you claiming it changes everything. It is deliberately ordered so that the cheapest disqualifying questions come first.
Appraisal is not a search for reasons to dismiss. The commonest failure among people who have just learned it is reflexive scepticism, which is indistinguishable in its effects from credulity: both end in practice unchanged by evidence. The aim is calibration — believing large, well-conducted, replicated findings about patients like yours, believing them proportionately less as each of those conditions weakens, and being able to say out loud which condition is doing the work in any given case.
Six calculators covering the numbers appraisal actually requires. All computation happens in your browser; nothing is transmitted or stored. Confidence intervals are labelled with the method used, because a proportion near 0 or 100 % gives nonsense under the textbook normal approximation and these use Wilson's method instead.
Three different families of tool get used interchangeably and should not be. Reporting guidelines (CONSORT, STROBE, PRISMA, STARD, TRIPOD, CHEERS) tell authors what to write down; a paper that omits an item may still be a good study, badly reported. Risk-of-bias tools (RoB 2, ROBINS-I, QUADAS-2, PROBAST, Newcastle–Ottawa) judge whether the study's result is likely to be distorted. Appraisal checklists (CASP) are teaching scaffolds for a reader forming a judgement.
The questions below are the working substance of each — enough to appraise with, phrased for the reader rather than the author. They are summaries for teaching, not reproductions: for a formal assessment, use the current official instrument from its own publisher.
QUADAS-2 assesses four domains for risk of bias, and the first three also for applicability to your own question.
| Column | Read it as |
|---|---|
| Participants / studies | How much evidence, and of what design |
| Relative effect | The transportable number |
| Anticipated absolute effects | The clinically meaningful number — check which baseline risk it assumes, and whether it is yours |
| Certainty | High / moderate / low / very low, per outcome, with footnotes naming the reason for every downgrade |
| Comments | Where the footnotes explaining the downgrades actually live — read these, not the letter grade |
Each of these is appraised for the method lesson it teaches, not as clinical guidance — several are chosen precisely because the clinical message is more complicated than the headline. Every figure, interval and p-value below was taken from the paper's own abstract retrieved from PubMed, and every author line and journal citation verified the same way. Where a number is not in the source, it is not here.
Read each card's lesson box last. The point of the sequence is that six real trials cover most of what goes wrong in reading: intention-to-treat, subgroups and interaction, a null primary with positive secondaries, external validity and harm, derivation versus validation, and the interval that a p-value hides.
What to notice in the methods. This is the archetype of the large simple trial: a broad eligibility criterion resting on clinician uncertainty, a trivially deliverable intervention, one hard outcome, and enough patients that a small true effect is detectable. The allocation mechanism deserves attention because it achieves concealment without any infrastructure — eight identical numbered packs in a box, so the enrolling clinician cannot know what the next patient will receive and therefore cannot steer sicker patients towards the drug.
Read the absolute numbers. The relative risk of 0.91 sounds modest; the absolute reduction is 16.0 % to 14.5 %, which is 1.5 percentage points and an NNT of about 68 to prevent one death, in a condition that kills one patient in six. Feed the counts into the treatment-effect calculator — it is preloaded with them — and confirm the published interval reproduces.
Where the interval sits. The upper bound is 0.97. That is a positive result whose weakest plausible version is a 3 % relative reduction in death — small, but on an outcome and at a baseline risk where small still matters. This is the pattern described in chapter 11 as "both ends clinically important, same direction", and it is why a modest p-value on a large trial with a hard outcome is stronger evidence than a striking p-value on a small one.
Why this subgroup analysis is believable when most are not. Apply the six-question test from chapter 14. The hypothesis was motivated in advance by the drug's mechanism — an antifibrinolytic given after fibrinolysis has resolved has no substrate to act on. The analysis is a formal test for interaction, not a hunt for significance within strata, and the interaction p-value is extreme. The gradient is monotonic across three ordered time bands. And critically, the authors report the interactions that were null in the same breath — blood pressure, GCS and injury type — which tells you the denominator of comparisons examined instead of leaving you to guess it.
Why it is nonetheless a subgroup finding. Randomisation guaranteed comparability within the trial as a whole; it did not guarantee it within a time stratum, because time to treatment is a post-randomisation characteristic correlated with injury pattern, transport distance and survival to enrolment. Patients treated after three hours are a different population from those treated within one, and not only in when they got the drug. The 3-hour signal is therefore the weakest of the three estimates despite its tidy confidence interval.
The appraisal problem this poses. A trial that misses both co-primary outcomes and hits two secondaries is the commonest genuinely difficult situation in clinical appraisal, and the temptation runs both ways: to dismiss the trial as negative, or to promote the secondary outcome as the real finding. Neither is right. What the trial licenses is a statement with the hierarchy preserved: 1:1:1 did not reduce mortality at 24 h or 30 days; it did reduce death from exsanguination and improve haemostasis, without a detectable safety cost from the extra plasma and platelets.
Three things make the secondary findings more than noise. They were pre-specified ancillary outcomes, not discovered afterwards. They are mechanistically coherent with the intervention and with each other — a ratio intended to correct coagulopathy reduced deaths from bleeding and produced more haemostasis. And they are consistent in direction with the null primary results, both of which favoured 1:1:1 numerically. A secondary outcome that pointed the opposite way to the primary would deserve far more suspicion.
Read the primary intervals rather than their p-values. The 24-hour difference of −4.2 % has an interval running from −9.6 % to +1.1 %. That is compatible with a large clinically decisive benefit and with a trivial harm — the "uninformative trial" pattern from chapter 11, and a much more honest description than "no difference". The trial is not evidence that ratio does not matter; it is evidence that this trial could not settle it at this sample size.
The trial's own premise is the lesson. The background section states that meta-analyses of small trials suggested tranexamic acid might decrease deaths from gastrointestinal bleeding — and a definitive trial of 12,009 patients found a risk ratio of 0.99. That is the small-study effect from chapter 19 playing out in real time, and it is the single best argument on this page for why a pooled estimate from small trials is a hypothesis rather than a conclusion.
Read this null result properly. The interval is 0.82 to 1.18 — tight, centred on no effect, and with both ends clinically unimportant. This is the pattern that genuinely licenses the phrase "no important difference", and it is quite different from ANDROMEDA-SHOCK's 0.55 to 1.02. A null result is only as informative as its interval is narrow, and here it is narrow because the trial was very large.
Why this is not a contradiction of CRASH-2. Different population (gastrointestinal bleeding, not trauma), different dose and duration (4 g over 24 h versus 2 g over 8), different primary outcome and different time frame. Fibrinolysis is central to the coagulopathy of major trauma in a way it is not to variceal or ulcer bleeding. The two trials are not in conflict; they are evidence that "does tranexamic acid help bleeding" is not a single question. Pooling them would produce an average describing no real patient.
The harm signal. Venous thromboembolism roughly doubled, on 48 events against 26, with an interval excluding 1. On a small number of events in a trial not primarily designed to measure thrombosis, this is a finding to take seriously without over-quantifying. Note the asymmetry from chapter 09: modified ITT dilutes benefit but also understates harm, since only patients who received the drug can be harmed by it.
Why the outcome definition is the most important line in the paper. ciTBI is defined by what happens to the child, not by what the scan shows. That is a deliberate choice, and it is what makes the rule usable: a tool built to detect any abnormality on CT would inevitably recommend scanning for findings that change nothing, while accepting more radiation in a population where radiation risk is the entire motivation. When a decision rule seems to tolerate "missing" injuries, check its outcome definition before objecting — the misses may be, by design, injuries that needed nothing.
Read the eligibility criterion as a hard boundary. The cohort is GCS 14–15 within 24 h of injury. The rule says nothing whatever about a child with GCS 13, and applying it there is not a cautious extension but a use outside validation. This is the single commonest bedside misapplication of any decision rule.
Read the negative predictive value with its prevalence. ciTBI occurred in 0.9 % of the whole cohort, so a tool that ruled out everybody would already achieve an NPV above 99 %. The NPV is impressive because it is paired with high sensitivity in a population where the outcome is genuinely rare — and because the confidence interval is reported. Note the 2-and-over rule's sensitivity: 61 of 63, with an interval down to 89.0 %. Two children with ciTBI were in the low-risk group, neither needing neurosurgery. A rule-out tool with a stated, small, characterised miss rate is a usable tool; one presenting itself as infallible is not.
Derivation and validation in one paper. The rules were derived in one part of the cohort and applied, fixed, to patients not used in building them. That is real validation, not a resubstitution estimate — but it is internal validation within one research network in North America, so transportability to other systems and case mixes is a separate question answered by separate work, and the rule's calibration in a lower-prevalence setting is not established by this paper.
What "p = 0.06" is being asked to carry here. The authors' conclusion — that the strategy did not reduce mortality — is the correct formal statement, because the trial did not meet its threshold. But read the interval alongside it. A hazard ratio of 0.75 with bounds of 0.55 and 1.02 is compatible with a 45 % relative reduction in death and with a 2 % relative increase. The absolute risk difference is −8.5 %, with an interval from −18.2 % to +1.2 %. Almost all of that interval lies in the region of substantial benefit. Summarising it as "capillary refill time made no difference" discards most of the information the trial produced, and the honest reading is: an inconclusive trial whose point estimate favours the intervention and which was not large enough to settle the question.
The threshold is a convention, not a finding. Had 3 more deaths occurred in the lactate arm, or 3 fewer in the perfusion arm, the result would have crossed p = 0.05 on a two-sided Fisher exact test of these counts, and the conclusion sentence in every subsequent citation would have been the opposite. Enter the counts — 74/212 against 92/212 — into the fragility calculator, which is preloaded with them, and see how few patients separate the two narratives. That fragility is the substance; the threshold is bookkeeping. (Fisher's exact test on the raw proportions gives p = 0.0906, not the 0.06 the trial reported from its time-to-event analysis — the two tests are asking slightly different questions of the same patients, which is itself worth noticing.)
And resist the opposite error. None of this makes the trial positive. A point estimate favouring the intervention on an inconclusive trial is a reason to want a larger trial, not a reason to change practice — and the SOFA difference of 1.00 point, with an upper bound of −0.02, sits on the edge of its own threshold and on a scale whose minimal clinically important difference is not established by this paper. Six other secondary outcomes did not differ. What the trial firmly establishes is that a bedside clinical target was not detectably worse than a laboratory one, which is itself a useful thing to know in a department without rapid lactate access.
Every number, confidence interval, p-value, author list, journal, volume and page reference on this tab was taken from the paper's own record retrieved from PubMed via the NCBI E-utilities interface on 26 September 2026, and nothing has been added from memory or from secondary summaries. Where a figure a reader might expect is absent — a number needed to treat, a fragility index, an absolute risk difference not reported by the authors — it is absent because the source does not state it, and the calculators on this site will compute it from the counts that the source does state.
These appraisals are teaching material. They describe what each trial reported and what method lesson it illustrates; they are not recommendations, and none of the six should be used to guide a treatment decision without reading the paper itself and the current guidance for the condition.
A results section is usually two or three figures and some prose explaining them. Recognising the figure and knowing what it conceals is a large fraction of appraisal, and it is quicker to learn than any of the arithmetic.
Each panel below says what you are looking at, then what it hides. Every figure here is computed, not drawn — from a published trial's own numbers where one exists, or from simulated data with a stated model where it does not, and each is labelled which. A funnel plot sketched by hand to look asymmetric would be invented data, so none of these were.
What it is. The map of every patient from screening to analysis, required by the CONSORT reporting guideline. It is the first figure in most trial reports and the one readers skip most reliably.
What to take from it. Two numbers per arm: how many were randomised, and how many were analysed. The difference between them, arm by arm, is the entire attrition story. CRASH-2 lost 36 of 10,096 and 48 of 10,115 — under half a per cent, and balanced — which is why its intention-to-treat analysis can be taken at face value.
What it is. A histogram bins the values and shows the shape of the distribution. A box plot compresses that shape to five numbers: the median line, the box spanning the interquartile range, whiskers out to the last point within 1.5 × IQR, and individual dots for anything beyond.
What to take from it. These are skewed data — a long right tail, as almost all emergency-department time data have. The mean (56.5 min) sits well above the median (38.5 min), and that gap is the skew. Reporting "mean 56.5, SD 44.7" would describe a patient experience almost nobody had: mean minus two SD is negative.
Paste your own series into the descriptive statistics calculator, which is preloaded with exactly these 40 numbers and draws both panels live.
What it is. The symmetric bell curve that a great deal of statistical machinery assumes. It is drawn here from the equation, not from data.
What to take from it. The fixed proportions are the whole reason the standard deviation is a useful summary: about 68 % of observations lie within one SD of the mean, 95 % within two, 99.7 % within three. This is also where laboratory reference ranges come from — the central 95 % of a healthy population — and therefore why one healthy person in twenty falls outside the range for any given test.
What it is. One row per study or per subgroup: a marker at the point estimate, a horizontal line for its confidence interval, and a vertical line at no effect. Marker area is proportional to weight, so the eye is drawn to the studies carrying the result. A pooled estimate, where there is one, is drawn as a diamond whose width is its interval.
What to take from it. This one uses CRASH-2's own counts for death due to bleeding, split by time to treatment. The reading is immediate in a way the numbers alone are not: the two early bands sit left of the line, the late band sits clearly right of it, and the intervals of the first and last do not overlap. That visual separation is the interaction described in chapter 14.
What it is. Every study in a meta-analysis plotted as effect size against precision, with the most precise studies at the top. Because small studies scatter more, an unbiased set of studies forms a symmetric inverted funnel around the true effect.
What to take from it. Both panels are the same simulated studies; the right-hand one simply has some small studies that found nothing removed, exactly as failure to publish would remove them. The result is a visible notch in the bottom corner on the null side, and a pooled estimate that would be pulled away from the truth. Asymmetry is a prompt to ask where the missing studies went.
What it is. The proportion still event-free over time, as a step function — each step down is an event, and censored patients (lost to follow-up, or still event-free when the study ended) leave the denominator without causing a step.
What to take from it. Read the numbers at risk underneath before anything else. They are the reason the right-hand end of any survival plot deserves suspicion: that is where the fewest patients remain and the uncertainty is widest, and it is exactly where curves are most often described as "separating".
What it is. Sensitivity against 1 − specificity, with one point for every possible cut-off of a continuous test. The diagonal is chance; the top-left corner is perfection. The area under the curve is the probability that a random patient with the disease scores higher than a random patient without it.
What to take from it. The two overlapping distributions on the left are what the curve is actually summarising. Where they separate, the curve bulges towards the corner; where they overlap, it collapses onto the diagonal. Each point on the curve is a threshold, so the curve is a menu of trade-offs rather than a single property of the test.
What it is. On the left, a conventional scatter of a new method against a reference with a fitted line and a correlation coefficient. On the right, the same data plotted as Bland–Altman: the difference between the two methods against their mean, with the mean difference (bias) and the limits of agreement.
What to take from it. The correlation is 0.990, which looks conclusive and answers the wrong question. Correlation measures whether two things move together; it is unchanged if one method reads a fixed amount high, or a fixed percentage high. The Bland–Altman plot asks whether they agree, and shows that the new method reads about 7 kg high, with an error that grows as the child gets bigger — a fan-shaped spread that is completely invisible on the left.
| Figure | Data | Source |
|---|---|---|
| CONSORT flow | Real | CRASH-2 randomisation and analysis counts, PMID 20554319 |
| Forest plot | Real | CRASH-2 death due to bleeding overall (PMID 20554319) and by time band (PMID 21439633). Intervals recomputed from the counts and confirmed to reproduce those published |
| Histogram / box plot | Simulated | 40 invented door-to-CT times, listed in full in the calculator |
| Normal curve | Idealised | Drawn from the equation; not data |
| Funnel plot | Simulated | 58 studies drawn around a true odds ratio of 0.80, with study size sampled continuously; the right panel removes small studies whose estimate fell on the null side |
| Kaplan–Meier | Simulated | 220 per arm, exponential survival; the crossing panel adds an early high-hazard fraction |
| ROC curve | Simulated | 300 with disease and 700 without, drawn from two overlapping normal distributions |
| Bland–Altman | Simulated | 90 paired measurements with a proportional bias deliberately built in |
Every figure is generated by a script held with the page source rather than drawn by hand, so each one is reproducible and can be checked against the numbers above. The simulated figures exist because no real dataset can be shown that is guaranteed to illustrate one specific artefact cleanly; they are labelled as simulated on the figure itself so they can never be mistaken for evidence about a real treatment.
A critical appraisal and evidence-based medicine curriculum written for emergency and acute care clinicians — the twenty-three chapters, six calculators, five appraisal checklists, six worked papers and eight annotated figures that between them cover what a clinician needs in order to read a paper and decide whether to act on it.
It is part of ResusDoc, a free set of UK emergency medicine tools built and maintained by Dr Nirmalya Hore, an NHS emergency physician. There are no adverts and no tracking beyond aggregate site analytics; nothing you type into a calculator leaves your device.
This is an educational resource about method. It is not clinical guidance. The six trials on the Worked papers tab were chosen because each teaches an appraisal trap cleanly, not because they represent current practice, and several were chosen precisely because their clinical message is more complicated than their headline. Nothing here should be used to make a treatment decision.
The checklists are teaching summaries, not the instruments. The questions under each are phrased for a reader forming a judgement, and they compress the substance of the named tools. For a formal risk-of-bias assessment or a submission, obtain the current official version of RoB 2, ROBINS-I, QUADAS-2, PROBAST, AMSTAR 2, AGREE II, CASP, CONSORT, STROBE, PRISMA, STARD, TRIPOD or CHEERS from its own publisher. Reporting guidelines are revised, and a summary written once will drift.
The calculators implement standard published methods — Wilson intervals for proportions, the log method for risk ratios, Woolf's method for odds ratios, Wald intervals for risk differences, a two-sided Fisher exact test for the fragility index, the normal-approximation formula for two-proportion sample size, and for the descriptive calculator the n−1 sample variance, quartiles by linear interpolation between order statistics, and the adjusted Fisher–Pearson skewness coefficient. They are intended for appraisal and teaching, and are not a substitute for a statistical package in the analysis of real data. Each one names the method it used alongside its output so that a result can be checked.
The figures are computed, not drawn. Every figure on the Figures tab is generated by a script kept with the page source, from either a published trial's own verified numbers or a stated simulation with a fixed seed, and each is labelled on the figure itself as real, simulated or idealised. None was sketched to look a certain way, because a plot drawn to illustrate an artefact is invented data. The provenance of all eight is tabulated at the foot of that tab.
What is deliberately not here. The mathematics behind any of the tests; the computation of a posterior distribution, as opposed to how to read one; machine-learning model evaluation; and formal instruction in conducting research rather than reading it. Each is a separate subject, and a page that covered all of them would teach none of them.
Every figure on the Worked papers tab was taken from the paper's own PubMed record, retrieved through the NCBI E-utilities interface on 26 September 2026 — including each author list, journal, volume, page range and PMID. No figure, byline or citation on that tab comes from memory or from a secondary summary, and where the source does not state a number it is not stated here either.
The two figures on the Figures tab that use real data — the CONSORT flow diagram and the forest plot — use CRASH-2's own counts, and the forest plot's confidence intervals were recomputed from those counts and confirmed to reproduce the published ones before the figure was used. The other six are simulated or idealised and say so on their face.
The curriculum chapters describe generally accepted appraisal method as taught in clinical epidemiology, and name the reporting guidelines and risk-of-bias tools by their published names. Where a chapter illustrates a point with a real trial, the figures are the verified ones from the Worked papers tab and no others.
Review status: internally reviewed only. There has been no external review of this page's content, statistical or educational. If you find an error — in a formula, an interval, a quoted figure or a description of a method — it is worth reporting, and corrections are made rather than defended.
The page is built to be usable in a departmental teaching session without preparation.
#t14 opens the subgroups chapter, #c-fagan the pre-test to post-test calculator, #k-quadas the diagnostic checklist, #f-funnel the funnel plot and #pp-proppr the PROPPR appraisal. Those links can go straight into a teaching email or a rota message.