Terug naar blog
Science & Safety 7 min2026-09-24

Tirzepatide vs Semaglutide: One Trial Randomized 751 Patients and Compared 12 Side Effects Head to Head

Almost every “tirzepatide vs semaglutide side effects” comparison splices together rates from different trials with different populations and durations — the same confound our dose-gradient piece found for dose. One trial in our corpus avoids it with both arms' registry counts: SURMOUNT-5 (NCT05822830) randomized 751 adults between the two drugs, titrated to each one's 15 mg or 2.4 mg maintenance dose, in the same protocol. Of 12 shared effects, tirzepatide's rate is lower for 7, tied for 1, and higher for 4 — including an 8.6%-vs-0.3% injection-site-reaction gap that a two-proportion test says is not noise, and two more gaps that only look that way. These are one trial's registry rows, not our pooled corpus estimates.

Dit artikel is nog niet in het Nederlands vertaald. Hieronder de Engelse versie.

Ask which GLP-1 drug has worse side effects, tirzepatide or semaglutide, and many of the comparisons you will find splice together two different trials — one drug's rate from one study, the other drug's rate from a different study, with a different population, a different duration and a different way of asking patients about symptoms. That is the same confound our dose-gradient piece found when comparing doses across trials, except here it is comparing drugs.

One trial in our corpus avoids that confound and gives both arms' counts. SURMOUNT-5 (NCT05822830) randomized 751 adults with obesity, or overweight with a weight-related comorbidity, to tirzepatide or semaglutide in the same protocol, over the same 72-week treatment period; 750 received at least one dose (374 and 376). It was open-label: participants and investigators knew which drug was given. (SURPASS-2, the type 2 diabetes head-to-head of tirzepatide against semaglutide 1 mg, is also in our corpus, but only through its journal abstract, which states tirzepatide's rates as dose ranges we withhold.) It is already in our data — until now, only for its published efficacy result (20.2% weight loss on tirzepatide against 13.7% on semaglutide, New England Journal of Medicine, 2025). Its posted ClinicalTrials.gov results record also carries arm-level adverse-event counts for twelve of the fifteen effects we track, and our registry collector stored all twenty-four rows — one tirzepatide row and one semaglutide row per effect — on 23 September 2026, verified against production the following day.

Read this as one trial, not our pooled estimate. Our pooled clinical rate for each effect lives at the data API and averages dozens of studies with different populations, doses and durations. This article is twenty-four registry rows from a single study, useful for exactly one reason: randomization inside it removes the population and duration confounds that make every cross-trial drug comparison on this site, or anywhere else, unreliable.

Both arms are a titration regimen, not a single dose

The trial's own results tables label both arms "or MTD" (maximum tolerated dose): tirzepatide "15 mg or MTD" (n=374 at risk) against semaglutide "2.4 mg or MTD" (n=376 at risk). Per the registry, tirzepatide started at 2.5 mg and rose by 2.5 mg every 4 weeks to 15 mg (or an MTD of 10 mg), and semaglutide started at 0.25 mg and rose every 4 weeks to 2.4 mg (or an MTD of 1.7 mg). Adverse events were counted from baseline to week 72, so every rate below includes the titration months as well as the maintenance dose. That arm label does not parse into our low/medium/high dose-tier system (both rows read "unspecified", the same MTD-labelling gap our dose-gradient article documented for arms like this one), so this comparison sits outside the predictor's tiered pooling entirely. It describes each drug's full titration-to-maintenance course, in this trial's population — not a starting dose on its own, and not a pure maintenance-dose comparison.

The twelve effects, ranked by the size of the gap

Every row is quoted directly from the trial's ClinicalTrials.gov results record, read on 23 September 2026 and re-read against the registry on 24 September 2026. "Acid reflux" is the registry term "Gastrooesophageal reflux disease"; "hair loss" is "Alopecia"; "reduced appetite" is "Decreased appetite". The registry lists a non-serious event only if it reached 5% in at least one arm, so rarer terms are not in the table:

Side effectTirzepatide 15 mg/MTD (n=374)Semaglutide 2.4 mg/MTD (n=376)Difference
Injection site reaction8.6% (32)0.3% (1)+8.3 pp
Vomiting15.0% (56)21.3% (80)−6.3 pp
Acid reflux6.1% (23)10.6% (40)−4.5 pp
Hair loss (alopecia)8.3% (31)6.1% (23)+2.2 pp
Fatigue10.4% (39)12.2% (46)−1.8 pp
Dizziness6.4% (24)4.8% (18)+1.6 pp
Constipation27.0% (101)28.5% (107)−1.5 pp
Nausea43.6% (163)44.4% (167)−0.8 pp
Reduced appetite4.5% (17)5.1% (19)−0.6 pp
Abdominal pain6.4% (24)6.9% (26)−0.5 pp
Headache7.2% (27)7.2% (27)0 pp
Diarrhoea23.5% (88)23.4% (88)+0.1 pp

Two example rows, verbatim: "In the '15 mg or MTD - Tirzepatide' arm (n=374 at risk), Injection site reaction occurred in 32 participants (8.6%)"; "In the '2.4 mg or MTD - Semaglutide' arm (n=376 at risk), Injection site reaction occurred in 1 participants (0.3%)." Every other cell in the table is the same registry's count divided by the same arm's at-risk denominator.

Read as a whole: tirzepatide's rate is lower for 7 of the 12 effects (abdominal pain, acid reflux, constipation, fatigue, nausea, reduced appetite, vomiting), tied for 1 (headache), and higher for 4 (diarrhoea, dizziness, hair loss, injection site reaction) — though one of those four, diarrhoea, is higher by a statistically meaningless 0.1 point.

What is real and what is noise

Twelve percentage-point gaps invite twelve stories. Most of them are not stories. Running a two-proportion z-test on each row (the standard large-sample approximation, no continuity correction) and applying a Bonferroni correction for testing twelve effects at once (critical z ≈ 2.87 for a family-wise α of 0.05, versus the usual 1.96 for a single test):

  • Injection site reaction (z ≈ 5.5): clears even the Bonferroni-corrected bar by a wide margin. This is not sampling noise.
  • Vomiting (z ≈ 2.2) and acid reflux (z ≈ 2.2): clear the conventional single-test line (|z| > 1.96) but not the corrected one. With twelve comparisons and no correction, roughly 0.6 "significant" results are expected by chance alone even if nothing genuinely differed — so these two are plausible but not the kind of finding we would stake a headline on alone. Both also sit in the same stomach-and-oesophagus family, where the registry lists related terms our taxonomy does not map (dyspepsia 22 against 28, eructation 37 against 29).
  • Hair loss (z ≈ 1.1) and dizziness (z ≈ 1.0): both look like real percentage-point gaps in the raw table and both are comfortably within sampling noise at these sample sizes. A 2-point gap on n≈375 per arm is not evidence of anything.
  • Every other row's gap is under 2 points and correspondingly weaker still.
  • We are publishing the z-values, not just the flag "significant" or not, because the table above will otherwise read as twelve findings when it supports at most three, and only one of those three without hedging.

    The injection-site-reaction gap

    32 participants of 374 against 1 of 376 is the largest, and the only unambiguous, gap in this comparison, from the same trial, population and treatment window. We do not put a ratio on it: with a single event in the semaglutide arm, any "N-fold" figure would rest on one person. We do not know why the gap exists, and the trial's own results record does not say. It could reflect the injection device or pen, how each product is injected, or a genuine difference in local tissue tolerability between the two molecules. Because the trial was open-label, differences in how readily a local reaction was reported cannot be ruled out either. We report the number; we are not going to guess at a mechanism the primary source does not give us.

    What this does not tell you

  • No placebo arm. SURMOUNT-5 is drug-versus-drug, not drug-versus-placebo, so nothing here separates either drug's true effect from a background symptom rate — unlike our constipation or headache pieces, which could read a placebo column because those trials had one.
  • One trial. A single randomized trial, however well it controls for population and duration, is still one sample. Our pooled, multi-study estimates for each effect — with their own stated n and source count — remain the better answer to "how common is this on this drug", and this article is not a substitute for them.
  • Open-label. Neither participants nor investigators were blinded to the drug, which can affect how symptoms are reported.
  • Whole course, not one dose. Each arm is a titration regimen counted over 72 weeks, so these rates mix the titration months with the maintenance dose and cannot be split between them.
  • No demographic breakdown. The registry's adverse-event module reports whole-arm counts; our stored rows carry no sex, age or BMI breakdown for either arm.
  • Multiple comparisons, restated. Twelve tests at the uncorrected 5% level carry about a 46% chance of at least one false "finding" by chance alone, even if nothing differed. Read the injection-site-reaction gap as established; read vomiting and acid reflux as suggestive; read hair loss and dizziness as noise dressed as a percentage.
  • If you are on a GLP-1 now

    This is educational content, not medical advice, and it does not tell you which drug is right for you — that depends on your own history, your prescriber's judgement, and factors this single trial's adverse-event table cannot capture. Nothing here should be used to start, stop or switch a medication without your prescriber.

    Every figure in this article was read from SURMOUNT-5's ClinicalTrials.gov results record on 23 September 2026, re-read against the registry and verified against our production store on 24 September 2026, the day this piece was first published. These twenty-four rows already count toward our corpus's pooled estimates for their effects — alongside dozens of other studies — at the data API, with the full eligibility methodology at magistra.health/en/methodology.

    Bekijk uw eigen cijfers

    Onze gratis voorspeller schat uw bijwerkingsrisico en gewichtsverloop, met bij elk cijfer het aantal vermelde percentages en afzonderlijke bronnen. Geen account nodig.

    Open de voorspeller

    Werkt u met de data zelf? De gedateerde momentopname achter deze cijfers is te koop als eenmalige aankoop, naast de gratis publieke API: Data & API.

    Bekijk uw persoonlijke GLP-1 bijwerkingsrisico

    Gratis Voorspeller