Back to blog
Science & Safety 4 min2026-08-21

Which GLP-1 Side Effects Have Enough Evidence to Calibrate a Statistical Model?

We track 15 GLP-1 side effects in our real-world evidence database. After a 31 August 2026 correction excluded seed rows from the eligible base, only 1 cleared our bar for a fully calibrated statistical model; as the corpus grew, 13 of the 15 clear it as of 30 September 2026. Here's which ones, why, and what's published instead for the rest.

Most GLP-1 content lists side-effect percentages without saying how many separate studies actually reported them, or whether a single well-publicised paper is doing all the work. We built Magistra's real-world evidence database specifically to answer that question for every effect we track — and the honest answer is: not evenly. Of the 15 side effects in our corpus, 12 now have enough stated, citable rate points to clear our bar for a fully calibrated statistical model — the Brier-score and Hosmer-Lemeshow diagnostics behind our model-health review — as of 18 September 2026. The three that did not were hair loss, injection-site reaction and emotional blunting. As of 08:45 UTC on 30 September 2026 it is 13: hair loss has since cleared the bar (17 eligible rates from 8 distinct sources), and injection-site reaction (9 eligible rates) and emotional blunting remain below it, per the same model-review endpoint. Effects below the bar still get a published pooled rate wherever the corpus has at least one eligible point; our companion piece has those numbers. This piece is about the higher calibration bar specifically, and we'd rather say plainly which effects clear it than blur the two together.

Correction, 18 September 2026. Until today this article's opening paragraph said in the present tense that "exactly 1" effect clears the calibration bar, and its two tables below are dated snapshots of 31 August 2026. The corpus has roughly trebled since (2,227 points across the 15 effects we publish — the corpusPoints figure on /api/data?q=overview), and 12 of the 15 effects clear the 10-eligible-point bar today — verified against our model-review endpoint (/api/data/review), whose modelFits array holds a fit for each of them, and cross-checked against the same response's review.issues entries, which flag exactly the other three as below the minimum. The dated tables below are left as the records of 31 August that they say they are; the live count is the endpoint linked beside each row, and that is the number to cite. This correction restates a count, not a rule: the 10-point bar and every eligibility check behind it are unchanged.

Correction, 31 August 2026. This article originally reported (as of 20 August 2026) that 5 effects cleared the calibration bar. On 31 August we found that the April-2026 pass that seeded our initial database had written rate rows restating our own static literature table under real trial and study URLs that no extraction ever read; URL-based eligibility screens cannot catch a row wearing a genuine URL, so 72 such rows sat inside the published eligible base. All are now labelled and excluded (details in our methodology's correction log), and every count below is restated over the cleaned base. Four of the original five effects fall below the 10-point bar as a result; the underlying rule is unchanged.

Looking for the actual pooled rate for a specific effect, not just whether it clears our bar? See every corpus-derived rate, with its confidence interval.

The rule, in full

A rate only counts toward a published estimate if it clears three checks, applied identically by our engine, our public API, and our offline audit script (documented in full at magistra.health/en/methodology):

  • Citable origin. The rate has to come from an external, linkable source — a clinical trial registry entry or a peer-reviewed paper. Self-referential citations (a rate pointing back at our own site) and news-aggregator search results (a headline about a source, not the source itself) don't count.
  • Incidence only, not spontaneous-report share. Our FAERS rows give a term's count as a share of the sum of that drug's thirty most-reported reaction-term counts — a share of term-mentions, not of reports and not of patients; FDA's own documentation calls these "reported events, not incidence rates." We track that separately and never blend it into a clinical rate.
  • One source, one vote. If a single paper states several rates, they collapse into one entry so it can't out-vote five independent studies. What we publish as "N distinct sources" is always post-collapse.
  • We then require at least 10 eligible rate points (post the checks above, pre-collapse) before we'll fit a statistically calibrated model for an effect — the diagnostic behind our model-health review. That bar is about the calibration check specifically, not about whether any number gets published at all: a simpler pooled rate (a weighted average with a confidence interval, its confidence grade reflecting how thin the base is) is published wherever at least one eligible point exists — see the companion piece for those.

    The 1 effect that clears the calibration bar (as of 31 August 2026)

    EffectStated clinical ratesDistinct sourcesCheck it live
    Nausea2919API

    Nausea is the gastrointestinal effect most consistently reported in GLP-1 trial literature, which is exactly why it clears the bar first — trial registries and peer-reviewed papers report it routinely, in numbers, as a primary or secondary adverse-event endpoint. The figure above is live-queryable at the link beside it; check our number rather than take our word for it.

    The 14 effects that don't yet clear the calibration bar (as of 31 August 2026)

    EffectEligible citable rate points
    Diarrhoea9
    Constipation9
    Reduced appetite9
    Vomiting7
    Headache5
    Hair loss3
    Dizziness3
    Abdominal pain1
    Acid reflux1
    Gallstones1
    Injection site reaction1
    Pancreatitis0
    Fatigue0
    Emotional blunting0

    Correction, 7 September 2026. Hair loss and dizziness moved from 0 to 3 eligible clinical rate points the same day this table was last checked, when a pinned fetch of SURMOUNT-1's posted adverse-event table recovered per-arm rates for both (3 stated rates, 1 distinct source each). Neither clears the 10-point calibration bar this article is about, so both stay in this table rather than joining nausea above — but both now have a published pooled rate rather than a literature fallback; see the companion piece for the numbers.

    One reading note: the counts in this table are citable rate points across both of our tracks — clinical sources and community reports that pass the same citability checks — because that combined count is what our 10-point calibration rule is applied to. The clinical-only subset for any effect (the measure used in the first table) is usually smaller, and is what the API link beside each effect reports as `rateBase.clinical`.

    This is a statement about our evidence base, not about the effects themselves. A low count here does not mean an effect is rare, mild, or doesn't happen — pancreatitis and gallstones are recognised serious risks of GLP-1 therapy regardless of how many trial abstracts state a numeric rate for them. It means that, as of today, we have not found enough stated, citable, numeric rates to run a fully calibrated statistical model for that specific effect — most commonly because the source material discusses the effect qualitatively (case reports, mechanism papers, warnings) rather than stating a comparable incidence percentage. We keep collecting daily; an effect moves from this table to the one above the day its 10th eligible, citable rate point lands. Eleven of these fourteen — including hair loss and dizziness as of the 7 September correction above — have a published pooled rate from the corpus today, each labelled with a confidence grade down to "very low" for a single- or few-source case — see the companion piece for the actual numbers. The other three — pancreatitis, fatigue and emotional blunting — have zero eligible clinical points, so both the predictor and the API fall back to a labelled static literature range (emotional blunting: no figure at all) instead of a corpus rate. You can check any of these effects yourself via the public API (for example, pancreatitis) — it returns whichever of the two currently applies.

    Why publish the gap at all

    Because a database that only shows you what it's confident about is easy to mistake for a database that has answered every question. We'd rather a researcher, journalist, or patient see "0 eligible points, not enough to calibrate" for pancreatitis than assume silence means we checked and found nothing to report — and rather publish a correction that shrinks our own headline (5 calibratable effects down to 1) than leave a wrong count standing. The full picture — what's well-evidenced, what isn't yet, and the exact rule that draws the line — is what makes the numbers above worth citing in the first place.

    See the complete, current picture — all 15 effects, live counts, confidence intervals, and every source — at magistra.health/en/data-api, or read the full eligibility methodology at magistra.health/en/methodology.

    See your own numbers

    Our free predictor estimates your side-effect risk and weight trajectory, with the stated rates and distinct sources shown behind every figure. No signup required.

    Open the predictor

    Working from the data itself? The dated snapshot behind these figures is available as a one-off purchase, alongside the free public API: Data & API.

    See your personal GLP-1 side-effect risk

    Free Predictor