Back to blog
Science & Safety 6 min2026-09-20

Our Best-Evidenced GLP-1 Side Effect Has a Wider Range Than One of Our Thinnest — Here's Why That's Not a Contradiction

Nausea, our most-evidenced side effect (135 stated rates, 29 distinct sources), publishes a 95% interval of 6–58%. Hair loss, one of our least-evidenced (8 rates, 3 sources), publishes 2–12% — visibly tighter. A reader skimming ranges alone would trust the narrow one more. Our own numbers say the opposite is closer to true, and the reason is a specific, checkable statistical mechanism, not a house quirk.

If you're comparing GLP-1 side-effect ranges across sources — ours or anyone's — here's a question worth asking before you trust the narrow one: narrow because the evidence is strong, or narrow because there isn't much evidence to disagree with itself? Our own two most extreme cases answer it directly, and the answer is not the one most readers would guess.

Nausea is the best-evidenced side effect in our database: 135 stated clinical rates from 29 distinct sources, source diversity graded `very_high` — as much independent trial and regulatory evidence as any effect we track. Its published 95% interval is 6–58%, around a pooled rate of 23.0%.

Hair loss is one of the thinnest: 8 stated rates from 3 distinct sources, graded `low`. Its interval is 2–12%, around a pooled rate of 4.5% — visibly, dramatically tighter than nausea's.

Read as a table (both figures computed the same way, live, by the same function, and checkable at the links):

EffectStated ratesDistinct sourcesConfidence gradePooled rate95% interval
Nausea13529very_high23.0%6–58%
Vomiting13327very_high10.3%3–33%
Hair loss83low4.5%2–12%
Injection-site reaction42low4.5%0–53%

Hair loss and injection-site reaction land on an identical 4.5% pooled rate. One publishes a range a fifth as wide as the other — 10 percentage points against 53 — on almost the same number of sources (3 versus 2). Three of the four rows defy what "more evidence narrows the range" would predict: our two deepest bases (29 and 27 sources) publish the second- and third-widest ranges here, and our 3-source row publishes the narrowest of the four. The fourth row, injection-site reaction, does land where the naive rule would put it — thinnest base, widest range — but for a reason the rule does not contain, which is what the next section is about. If interval width were the signal to read, this table would be actively misleading.

What's actually happening

Our interval isn't sampling error alone — it's a random-effects interval (documented in full at our methodology, §2.5), which adds a between-study term, τ², that grows with how much the individual source rates disagree with each other, not with how many sources there are. Two things follow, and both are visible in the table above:

  • Few sources that happen to agree produce a narrow interval — and that narrowness is fragile, not earned. Hair loss's 3 sources currently sit close together, so τ² is small and the interval looks tight. This is not hypothetical for this exact effect: in our hair-loss piece, the interval was 4–7% on 1 source, then 4–7% on 2, then widened to 2–12% the day a third source landed with a rate well below the first two. A fourth source could move it again, in either direction, by more than a well-evidenced effect's interval moves in a month. Three sources is not "settled" — it's "not yet contradicted."
  • Many sources that genuinely disagree produce a wide interval — and that width is the evidence working correctly, not failing. Nausea's 29 sources span trials in different populations, drugs, doses and durations; real trials of a common, dose-related, multi-week side effect do not converge on one number, and 29 independent readings saying so is a stronger, truer answer than a falsely narrow one would be. Averaging that disagreement away would hide something real about how much nausea risk actually varies by who you are and what you're on — which is closer to the point of a personalised predictor than a single headline percentage ever was.
  • Injection-site reaction makes the same point starkly with almost no data at all: its 2 sources post 14.3% and 4.4% — a threefold gap — and it is that disagreement, not the thinness on its own, that opens its interval to 0–53%, more than half the range a percentage can occupy. This is the row the naive rule gets right by accident: a thin base does not have to produce a wide range, as hair loss's three closely-agreeing sources show one line above it. That 0–53% is an honest "we don't know within a factor of ten yet" — which a single, cleaner-looking number would have hidden.

    The number to actually read

    Not the width of the range — the source-diversity grade and the source count published beside it. Our API states both on every pooled estimate: `sourceDiversity` (`very_low` through `very_high`) and `distinctSources`, plus an `intervalBasis` field that says outright whether an interval reflects one source's sampling noise alone (`sampling_only_single_source`) or genuine between-study spread. A `low`-diversity effect with a narrow-looking range is exactly the case those fields exist to flag, and it is why we publish them next to the percentage rather than leaving the percentage to speak for itself. This generalises past our own site: any side-effect number quoted without its source count invites exactly this misread, in either direction.

    Two companion pieces go deeper on the pieces either side of this one: which effects clear the bar for a fully calibrated model covers source count thresholds directly, and what a pooled rate is made of shows that even a well-evidenced pool can be carried disproportionately by one drug's rows. This piece is about the number people are likeliest to misread without either: the range itself.

    Reproduce it

    Every figure above was computed on 20 September 2026 by the same pooling function that runs live on magistra.health (`pooledClinicalEstimate`, documented in full), run directly against our corpus rather than argued from a summary — cross-checked against the live-verified count published in yesterday's composition piece (1,142 stated rates, 31 distinct studies, unchanged since). Read any effect's current numbers, including `sourceDiversity` and `intervalBasis`, at `GET /api/data?q=effect&id=` — the corpus grows daily, so an interval quoted here can and will move.

    The full dataset — 15 effects, both tracks, every source with its URL, CC BY 4.0 — is at magistra.health/en/data-api.

    See your own numbers

    Our free predictor estimates your side-effect risk and weight trajectory, with the stated rates and distinct sources shown behind every figure. No signup required.

    Open the predictor

    Working from the data itself? The dated snapshot behind these figures is available as a one-off purchase, alongside the free public API: Data & API.

    See your personal GLP-1 side-effect risk

    Free Predictor