Does a Higher GLP-1 Dose Mean Worse Side Effects? We Compared 78 Dose Arms Inside the Same Trials
We measured the dose gradient the only fair way posted trial results allow: between arms of the same trial, on the same drug. Across 78 comparisons from 8 trials, going from a trial's lowest dose arm to its highest adds a median of 2.7 percentage points to gut effects and 0.8 points to everything else — far less than the framing implies. It also shows that the 23%-to-42% high-dose jump on our own predictor is mostly a change of drug mix, not of dose. And of our 1,142 eligible rates, zero describe a starting dose.
Ask an AI "will my side effects get worse when I go up a dose?" and you will get a confident yes. The drug labels encourage it: they report one whole-trial incidence per effect for a maintenance-dose arm, prescribers titrate slowly precisely because of side effects, and every patient forum treats a dose increase as a week to brace for.
We have 1,142 eligible stated rates from 31 distinct studies, so we went and measured it. The answer is yes for the gut, and much smaller than the framing suggests — a median of +2.7 percentage points from a trial's lowest dose arm to its highest. For everything that is not a gut effect, +0.8 points, which is close to nothing.
Along the way we found that the biggest dose-related number on our own site is mostly not about dose. That is the more useful half of this article.
Why you have to compare inside one trial
The obvious way to do this is the wrong way: pool every high-dose rate in the corpus, pool every low-dose rate, compare the two. Those two pools differ in population, trial duration, how the adverse event was solicited, and — as we'll see below — which molecule was being given. Any gap you find is all of that at once.
The fair comparison is between arms of the same trial: one protocol, one population, one adverse-event definition, one results table posted by the same sponsor on the same day. ClinicalTrials.gov posts an at-risk denominator and an event count for every arm separately, so the comparison is available wherever a trial ran more than one dose.
Four restrictions, each removing a specific confound rather than tuning the answer:
The mistake we made first, because it is the one everyone makes
Our first pass keyed each comparison on trial and effect. That produced a clean, satisfying dose gradient — nausea +6.4 points, vomiting +8.3, constipation +4.7.
Then we read the arm labels. A trial's arms are not all the same molecule. ACHIEVE-3 (NCT06045221) posts oral semaglutide 7 mg and 14 mg arms beside an orforglipron 36 mg arm; SURPASS-SWITCH (NCT05564039) posts dulaglutide 4.5 mg beside tirzepatide 15 mg. Keyed on trial and effect alone, both of those became "4.5 mg → 15 mg, +6.3 points of nausea" — a drug difference wearing a dose label, and both landed in the largest-mover rows.
Requiring the same drug within a comparison took constipation from +4.7 points to +1.2, abdominal pain from +1.7 to +0.4, and the gut-effect median from +3.7 to +2.7. The confound was inflating the answer by roughly a third.
What 78 same-trial, same-drug comparisons show
78 effect × trial × drug comparisons across 8 trials (orforglipron 49, tirzepatide 16, semaglutide 13), read from our production store on 22 September 2026. Each row is the median change from a trial's lowest qualifying dose arm to its highest.
| Effect | Comparisons | Median change | Median ratio | Rises | Falls |
|---|---|---|---|---|---|
| Vomiting | 8 | +5.8 pp | 1.50× | 6 | 2 |
| Nausea | 8 | +5.5 pp | 1.45× | 8 | 0 |
| Diarrhoea | 8 | +4.2 pp | 1.30× | 7 | 1 |
| Reduced appetite | 7 | +3.7 pp | 1.76× | 5 | 2 |
| Fatigue | 6 | +1.8 pp | 1.62× | 5 | 1 |
| Injection site reaction | 1 | +1.7 pp | 1.59× | 1 | 0 |
| Constipation | 9 | +1.2 pp | 1.16× | 6 | 3 |
| Headache | 6 | +1.1 pp | 1.22× | 4 | 2 |
| Dizziness | 5 | +1.1 pp | 1.19× | 3 | 2 |
| Hair loss | 2 | +1.1 pp | 1.13× | 2 | 0 |
| Acid reflux | 6 | +0.9 pp | 1.25× | 6 | 0 |
| Abdominal pain | 7 | +0.4 pp | 1.08× | 5 | 2 |
| Pancreatitis | 2 | +0.1 pp | — | 1 | 1 |
| Gallstones | 3 | −0.1 pp | 0.88× | 1 | 2 |
Split by mechanism:
Nausea rises in all 8 of its comparisons, the largest all-rising set in the table — acid reflux is 6 of 6 and hair loss 2 of 2, on much thinner samples — which is a stronger signal than its modest median move suggests. Gallstones is the only effect whose median goes the other way, and with 3 comparisons and a −0.1 point median that is a null, not a protective effect.
The practical reading: going from the bottom of a trial's dose range to the top adds a few points of gut trouble and leaves your headache risk roughly where it was. Nausea at 25.4% on tirzepatide 5 mg is 31.9% on 15 mg in SURMOUNT-1 (NCT04184622, n=630 per arm) — a tripling of dose for a 6.5-point move.
The number on our own site that is mostly not about dose
As of today our predictor publishes, for the high dose tier, a rate pooled only from trial arms explicitly tagged to the top rung of an approved maintenance ladder. For nausea that is 42%, against 23% pooled across all doses. A 19-point gap, on a page about dose tiers, invites exactly one reading.
That reading would be wrong, and the table above is how we know. Within a trial and within a drug, the top-to-bottom dose move on nausea is 5.5 points, not 19.
Here is what is actually happening. Our eligible base is 876 of 1,142 rates from orforglipron — an oral GLP-1 (approved as Foundayo, April 2026) whose dose-ranging trials post a great many arms. The all-dose nausea pool is 103 of 135 rates orforglipron, and orforglipron's trials report lower nausea than tirzepatide's. The high-tier pool contains zero orforglipron rates: the maintenance ladder we map doses onto only exists for semaglutide, tirzepatide and dulaglutide, so nausea's high-tier figure is 4 semaglutide rates and 2 tirzepatide rates and nothing else.
So "23% → 42%" is substantially a change of drug mix, not a change of dose. Both figures are correct as computed, both carry their n and their source count, and the comparison between them means less than it looks like it means. We have filed that as an open disclosure question rather than quietly adjusting the wording, because it changes how a live figure is explained on fourteen effects.
The dose nobody reports
There is a second, larger gap underneath all of this.
Of our 1,142 eligible clinical and regulatory rates, the number we can place on the starting rung of an approved maintenance ladder — semaglutide below 1.0 mg, tirzepatide below 5 mg, dulaglutide below 1.5 mg — is zero. Not few. None. The census is 62 at the top rung, 37 in the middle, 1,043 that cannot be placed on a ladder at all (mostly orforglipron and retatrutide — approved and investigational respectively, but neither mapped onto a tier ladder in our system — plus titrate-to-maximum-tolerated-dose arms and arm labels that state no dose), and 0 at the bottom.
This is not a gap in our collection. It is what the trials do. Registries report maintenance-dose or maximum-tolerated-dose arms; the starting dose exists only as the first few weeks of a titration path inside those arms, and nobody posts an adverse-event table for it. So the question most people actually have on day one — how common is nausea in my first month at 0.25 mg? — cannot be answered from posted trial results at all, by us or by anyone.
We say so on the tool rather than inventing a number: at the low tier the predictor discloses that no corpus-derived rate exists for that dose tier and that the adjustment it applies comes from a static literature reference table, not from this corpus.
Honest limitations
Reproduce it
The measurement is one committed script, measure-dose-gradients.mjs, which imports the same eligibility predicate the live site uses so it cannot drift from what we publish. We publish it so you can read exactly what was computed, not so you can run it: like everything else in that mirror it is a snapshot that expects our application tree and our store, so it is the audit trail rather than something you can execute yourself. We ran it twice for this article on 22 September 2026 — once against the repo's committed corpus and once against production with `--kv` — and the two returned identical figures on every row above. Its `--no-filters` mode, which we also ran, prints what the four restrictions cost.
Every arm, denominator and event count above is in the ClinicalTrials.gov record linked beside it. The pooled figures are from our live endpoint, and the full dataset — 15 effects, both tracks, every source with its URL and stated sample size, CC BY 4.0 — is at magistra.health/en/data-api, with the eligibility methodology at magistra.health/en/methodology.
All figures read from production on 22 September 2026, and they move with each collection run — verify at the endpoint before citing.
Bekijk uw eigen cijfers
Onze gratis voorspeller schat uw bijwerkingsrisico en gewichtsverloop, met bij elk cijfer het aantal vermelde percentages en afzonderlijke bronnen. Geen account nodig.
Open de voorspellerWerkt u met de data zelf? De gedateerde momentopname achter deze cijfers is te koop als eenmalige aankoop, naast de gratis publieke API: Data & API.