Our Best-Evidenced GLP-1 Side Effect Has a Wider Range Than One of Our Thinnest — Here's Why That's Not a Contradiction
Nausea, our most-evidenced side effect (135 stated rates, 29 distinct sources), publishes a 95% interval of 6–58%. Hair loss, one of our least-evidenced (8 rates, 3 sources), publishes 2–12% — visibly tighter. A reader skimming ranges alone would trust the narrow one more. Our own numbers say the opposite is closer to true, and the reason is a specific, checkable statistical mechanism, not a house quirk.
If you're comparing GLP-1 side-effect ranges across sources — ours or anyone's — here's a question worth asking before you trust the narrow one: narrow because the evidence is strong, or narrow because there isn't much evidence to disagree with itself? Our own two most extreme cases answer it directly, and the answer is not the one most readers would guess.
Nausea is the best-evidenced side effect in our database: 135 stated clinical rates from 29 distinct sources, source diversity graded `very_high` — as much independent trial and regulatory evidence as any effect we track. Its published 95% interval is 6–58%, around a pooled rate of 23.0%.
Hair loss is one of the thinnest: 8 stated rates from 3 distinct sources, graded `low`. Its interval is 2–12%, around a pooled rate of 4.5% — visibly, dramatically tighter than nausea's.
Read as a table (both figures computed the same way, live, by the same function, and checkable at the links):
| Effect | Stated rates | Distinct sources | Confidence grade | Pooled rate | 95% interval |
|---|---|---|---|---|---|
| Nausea | 135 | 29 | very_high | 23.0% | 6–58% |
| Vomiting | 133 | 27 | very_high | 10.3% | 3–33% |
| Hair loss | 8 | 3 | low | 4.5% | 2–12% |
| Injection-site reaction | 4 | 2 | low | 4.5% | 0–53% |
Hair loss and injection-site reaction land on an identical 4.5% pooled rate. One publishes a range a fifth as wide as the other — 10 percentage points against 53 — on almost the same number of sources (3 versus 2). Three of the four rows defy what "more evidence narrows the range" would predict: our two deepest bases (29 and 27 sources) publish the second- and third-widest ranges here, and our 3-source row publishes the narrowest of the four. The fourth row, injection-site reaction, does land where the naive rule would put it — thinnest base, widest range — but for a reason the rule does not contain, which is what the next section is about. If interval width were the signal to read, this table would be actively misleading.
What's actually happening
Our interval isn't sampling error alone — it's a random-effects interval (documented in full at our methodology, §2.5), which adds a between-study term, τ², that grows with how much the individual source rates disagree with each other, not with how many sources there are. Two things follow, and both are visible in the table above:
Injection-site reaction makes the same point starkly with almost no data at all: its 2 sources post 14.3% and 4.4% — a threefold gap — and it is that disagreement, not the thinness on its own, that opens its interval to 0–53%, more than half the range a percentage can occupy. This is the row the naive rule gets right by accident: a thin base does not have to produce a wide range, as hair loss's three closely-agreeing sources show one line above it. That 0–53% is an honest "we don't know within a factor of ten yet" — which a single, cleaner-looking number would have hidden.
The number to actually read
Not the width of the range — the source-diversity grade and the source count published beside it. Our API states both on every pooled estimate: `sourceDiversity` (`very_low` through `very_high`) and `distinctSources`, plus an `intervalBasis` field that says outright whether an interval reflects one source's sampling noise alone (`sampling_only_single_source`) or genuine between-study spread. A `low`-diversity effect with a narrow-looking range is exactly the case those fields exist to flag, and it is why we publish them next to the percentage rather than leaving the percentage to speak for itself. This generalises past our own site: any side-effect number quoted without its source count invites exactly this misread, in either direction.
Two companion pieces go deeper on the pieces either side of this one: which effects clear the bar for a fully calibrated model covers source count thresholds directly, and what a pooled rate is made of shows that even a well-evidenced pool can be carried disproportionately by one drug's rows. This piece is about the number people are likeliest to misread without either: the range itself.
Reproduce it
Every figure above was computed on 20 September 2026 by the same pooling function that runs live on magistra.health (`pooledClinicalEstimate`, documented in full), run directly against our corpus rather than argued from a summary — cross-checked against the live-verified count published in yesterday's composition piece (1,142 stated rates, 31 distinct studies, unchanged since). Read any effect's current numbers, including `sourceDiversity` and `intervalBasis`, at `GET /api/data?q=effect&id=
The full dataset — 15 effects, both tracks, every source with its URL, CC BY 4.0 — is at magistra.health/en/data-api.
Bekijk uw eigen cijfers
Onze gratis voorspeller schat uw bijwerkingsrisico en gewichtsverloop, met bij elk cijfer het aantal vermelde percentages en afzonderlijke bronnen. Geen account nodig.
Open de voorspellerWerkt u met de data zelf? De gedateerde momentopname achter deze cijfers is te koop als eenmalige aankoop, naast de gratis publieke API: Data & API.