Terug naar blog
Science & Safety 6 min2026-09-20

Our Best-Evidenced GLP-1 Side Effect Has a Wider Range Than One of Our Thinnest — Here's Why That's Not a Contradiction

Nausea, our most-evidenced side effect (135 stated rates, 29 distinct sources), publishes a 95% interval of 6–58%. Hair loss, one of our least-evidenced (8 rates, 3 sources), publishes 2–12% — visibly tighter. A reader skimming ranges alone would trust the narrow one more. Our own numbers say the opposite is closer to true, and the reason is a specific, checkable statistical mechanism, not a house quirk.

Dit artikel is nog niet in het Nederlands vertaald. Hieronder de Engelse versie.

If you're comparing GLP-1 side-effect ranges across sources — ours or anyone's — here's a question worth asking before you trust the narrow one: narrow because the evidence is strong, or narrow because there isn't much evidence to disagree with itself? Our own two most extreme cases answer it directly, and the answer is not the one most readers would guess.

Nausea is the best-evidenced side effect in our database: 135 stated clinical rates from 29 distinct sources, source diversity graded `very_high` — as much independent trial and regulatory evidence as any effect we track. Its published 95% interval is 6–58%, around a pooled rate of 23.0%.

Hair loss is one of the thinnest: 8 stated rates from 3 distinct sources, graded `low`. Its interval is 2–12%, around a pooled rate of 4.5% — visibly, dramatically tighter than nausea's.

Read as a table (both figures computed the same way, live, by the same function, and checkable at the links):

EffectStated ratesDistinct sourcesConfidence gradePooled rate95% interval
Nausea13529very_high23.0%6–58%
Vomiting13327very_high10.3%3–33%
Hair loss83low4.5%2–12%
Injection-site reaction42low4.5%0–53%

Hair loss and injection-site reaction land on an identical 4.5% pooled rate. One publishes a range a fifth as wide as the other — 10 percentage points against 53 — on almost the same number of sources (3 versus 2). Three of the four rows defy what "more evidence narrows the range" would predict: our two deepest bases (29 and 27 sources) publish the second- and third-widest ranges here, and our 3-source row publishes the narrowest of the four. The fourth row, injection-site reaction, does land where the naive rule would put it — thinnest base, widest range — but for a reason the rule does not contain, which is what the next section is about. If interval width were the signal to read, this table would be actively misleading.

What's actually happening

Our interval isn't sampling error alone — it's a random-effects interval (documented in full at our methodology, §2.5), which adds a between-study term, τ², that grows with how much the individual source rates disagree with each other, not with how many sources there are. Two things follow, and both are visible in the table above:

  • Few sources that happen to agree produce a narrow interval — and that narrowness is fragile, not earned. Hair loss's 3 sources currently sit close together, so τ² is small and the interval looks tight. This is not hypothetical for this exact effect: in our hair-loss piece, the interval was 4–7% on 1 source, then 4–7% on 2, then widened to 2–12% the day a third source landed with a rate well below the first two. A fourth source could move it again, in either direction, by more than a well-evidenced effect's interval moves in a month. Three sources is not "settled" — it's "not yet contradicted."
  • Many sources that genuinely disagree produce a wide interval — and that width is the evidence working correctly, not failing. Nausea's 29 sources span trials in different populations, drugs, doses and durations; real trials of a common, dose-related, multi-week side effect do not converge on one number, and 29 independent readings saying so is a stronger, truer answer than a falsely narrow one would be. Averaging that disagreement away would hide something real about how much nausea risk actually varies by who you are and what you're on — which is closer to the point of a personalised predictor than a single headline percentage ever was.
  • Injection-site reaction makes the same point starkly with almost no data at all: its 2 sources post 14.3% and 4.4% — a threefold gap — and it is that disagreement, not the thinness on its own, that opens its interval to 0–53%, more than half the range a percentage can occupy. This is the row the naive rule gets right by accident: a thin base does not have to produce a wide range, as hair loss's three closely-agreeing sources show one line above it. That 0–53% is an honest "we don't know within a factor of ten yet" — which a single, cleaner-looking number would have hidden.

    The number to actually read

    Not the width of the range — the source-diversity grade and the source count published beside it. Our API states both on every pooled estimate: `sourceDiversity` (`very_low` through `very_high`) and `distinctSources`, plus an `intervalBasis` field that says outright whether an interval reflects one source's sampling noise alone (`sampling_only_single_source`) or genuine between-study spread. A `low`-diversity effect with a narrow-looking range is exactly the case those fields exist to flag, and it is why we publish them next to the percentage rather than leaving the percentage to speak for itself. This generalises past our own site: any side-effect number quoted without its source count invites exactly this misread, in either direction.

    Two companion pieces go deeper on the pieces either side of this one: which effects clear the bar for a fully calibrated model covers source count thresholds directly, and what a pooled rate is made of shows that even a well-evidenced pool can be carried disproportionately by one drug's rows. This piece is about the number people are likeliest to misread without either: the range itself.

    Reproduce it

    Every figure above was computed on 20 September 2026 by the same pooling function that runs live on magistra.health (`pooledClinicalEstimate`, documented in full), run directly against our corpus rather than argued from a summary — cross-checked against the live-verified count published in yesterday's composition piece (1,142 stated rates, 31 distinct studies, unchanged since). Read any effect's current numbers, including `sourceDiversity` and `intervalBasis`, at `GET /api/data?q=effect&id=` — the corpus grows daily, so an interval quoted here can and will move.

    The full dataset — 15 effects, both tracks, every source with its URL, CC BY 4.0 — is at magistra.health/en/data-api.

    Bekijk uw eigen cijfers

    Onze gratis voorspeller schat uw bijwerkingsrisico en gewichtsverloop, met bij elk cijfer het aantal vermelde percentages en afzonderlijke bronnen. Geen account nodig.

    Open de voorspeller

    Werkt u met de data zelf? De gedateerde momentopname achter deze cijfers is te koop als eenmalige aankoop, naast de gratis publieke API: Data & API.

    Bekijk uw persoonlijke GLP-1 bijwerkingsrisico

    Gratis Voorspeller