Trial data comparison (revision 62)
Old revision·16:45, 4 Feb 2026·CategoryBot
| Trial data comparisonInteractive reference tool | |
|---|---|
Effect sizes from separate trials. Placing them side by side is not a head-to-head comparison. | |
| Rows | Major randomised trials of incretin agonists |
| Columns | Population, n, duration, comparator, endpoint, result |
| Caveats | |
| Head-to-head? | Only where the comparator column says so |
| Estimand | Stated per row where known |
| Populations | Differ substantially between trials |
| Reference tool infobox · conventions | |
The trial data comparison tabulates the principal published results of the major randomised trials of GLP-1 receptor agonists and related compounds, so that a reader arriving at one figure can see the others alongside it. Each row is attributed to its trial, arm, duration and comparator.
The table is easy to misread in one specific way, and the article exists partly to prevent it. Two figures from two trials are not a comparison of two drugs. Trial populations differ in baseline weight, glycaemic status, diabetes duration, background therapy and geography; durations differ; and the statistical estimand used to summarise a weight-change endpoint differs between publications and sometimes within a single one.[1] A larger number in this table means a larger result in that trial, and nothing more.
Where a genuine head-to-head randomised comparison exists — SURPASS-2 compared tirzepatide with semaglutide 1 mg directly — the comparator column says so, and only those rows support a statement that one agent outperformed another.[2]
The comparison
[edit]Filter by compound, trial programme or endpoint type. The bar is drawn to the magnitude of the primary result and is scaled within its endpoint type only.
Why the estimand matters
[edit]A weight-change endpoint can be summarised in at least two defensible ways, and the ICH E9(R1) addendum names them.[1]
The treatment-policy estimand asks what happened to everyone randomised, including those who stopped the drug and those who started another one. It answers "what does prescribing this achieve", and it is the more conservative figure.
The trial-product estimand asks what happened while participants were taking the drug as intended. It answers "what does the molecule do", and it is the larger figure.
The gap between them is not small. In the STEP programme the two estimands for the same trial differed by roughly one to two percentage points of body weight, which is a difference of the same order as the gap between some of the drugs being compared.[3] A table that mixes estimands between rows therefore manufactures apparent differences between compounds that are artefacts of the analysis choice.
| Estimand | Question answered | Effect on the figure |
|---|---|---|
| Treatment policy | What prescribing achieves | Smaller; includes discontinuations |
| Trial product | What the molecule does on treatment | Larger; censors discontinuation |
| Not stated | — | Unusable for comparison |
The SURMOUNT programme's publications state their estimand explicitly, which is why the tirzepatide rows in the table above can be compared with each other with more confidence than the cross-programme rows can.[4]
What differs between these trials
[edit]Five differences are large enough to dominate any cross-trial reading.
Glycaemic status. Trials in people with type 2 diabetes consistently report smaller weight reductions than trials in people without it, at the same dose of the same drug. SURPASS and SURMOUNT are not comparable on weight for this reason alone.
Baseline weight. A percentage reduction from a higher baseline is a larger absolute loss. Reporting one and not the other changes the apparent ordering.
Duration. Weight-loss curves in this class had not fully plateaued at 68 weeks in several trials, so a 40-week and a 72-week figure are not measuring the same thing.
Background therapy. Trials differ in whether metformin, insulin or a sulfonylurea was permitted, which affects both glycaemic and weight endpoints and the hypoglycaemia rate.
Comparator. Placebo-controlled and active-controlled trials answer different questions. A placebo-adjusted difference and an absolute change are frequently confused; see Placebo-adjusted effect.
Outcome trials are a different kind of evidence
[edit]Three of the rows report hazard ratios for clinical events rather than changes in a measurement. These are the strongest evidence in the table and the least comparable to the rest of it: an event-driven trial reports a relative risk over years in a population selected for risk, and a hazard ratio cannot be placed alongside a percentage weight change in any meaningful ordering.
For those rows the useful companion figure is the number needed to treat, which converts a relative effect into an absolute one over a stated horizon — and which is meaningless without that horizon.[5]
See also
- STEP trial programme
- SURMOUNT trial programme
- SURPASS trial programme
- SELECT trial
- Intention-to-treat analysis
- Placebo-adjusted effect
- Number needed to treat
References
- ^ a b International Council for Harmonisation, E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials (2019).
- ^ Frías JP, Davies MJ, Rosenstock J, et al. "Tirzepatide versus Semaglutide Once Weekly in Patients with Type 2 Diabetes." New England Journal of Medicine 385(6):503–515 (2021). DOI:10.1056/NEJMoa2107519. PMID 34170647.
- ^ Wilding JPH, Batterham RL, Calanna S, et al. "Once-Weekly Semaglutide in Adults with Overweight or Obesity." New England Journal of Medicine 384(11):989–1002 (2021). DOI:10.1056/NEJMoa2032183. PMID 33567185.
- ^ Jastreboff AM, Aronne LJ, Ahmad NN, et al. "Tirzepatide Once Weekly for the Treatment of Obesity." New England Journal of Medicine 387(3):205–216 (2022). DOI:10.1056/NEJMoa2206038. PMID 35658024.
- ^ Lincoff AM, Brown-Frandsen K, Colhoun HM, et al. "Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes." New England Journal of Medicine 389(24):2221–2232 (2023). DOI:10.1056/NEJMoa2307563. PMID 37952131.