Which GLP-1 Side Effects Have Enough Evidence to Calibrate a Statistical Model?
We track 15 GLP-1 side effects in our real-world evidence database. Only 5 currently clear our bar for a fully calibrated statistical model. Here's exactly which ones, why, and what's published instead for the other 10.
Most GLP-1 content lists side-effect percentages without saying how many separate studies actually reported them, or whether a single well-publicised paper is doing all the work. We built Magistra's real-world evidence database specifically to answer that question for every effect we track — and the honest answer is: not evenly. Of the 15 side effects in our corpus, 5 currently have enough independently-sourced, citable evidence to clear our bar for a fully calibrated statistical model — the Brier-score and Hosmer-Lemeshow diagnostics behind our model-health review. The other 10 don't clear that bar yet, but they still get a published pooled rate wherever the corpus has at least one eligible point; our companion piece has those numbers. This piece is about the higher calibration bar specifically, and we'd rather say plainly which effects clear it than blur the two together.
Looking for the actual pooled rate for a specific effect, not just whether it clears our bar? See [every corpus-derived rate, with its confidence interval](/en/blog/glp1-side-effect-rates-corpus-derived).
The rule, in full
A rate only counts toward a published estimate if it clears three checks, applied identically by our engine, our public API, and our offline audit script (documented in full at magistra.health/en/methodology):
We then require at least 10 eligible rate points (post the checks above, pre-collapse) before we'll fit a statistically calibrated model for an effect — the diagnostic behind our model-health review. That bar is about the calibration check specifically, not about whether any number gets published at all: a simpler pooled rate (a weighted average with a confidence interval, its confidence grade reflecting how thin the base is) is published wherever at least one eligible point exists — see the companion piece for those.
The 5 effects that clear the calibration bar (as of 20 August 2026)
| Effect | Independent clinical rates | Distinct sources | Check it live |
|---|---|---|---|
| Nausea | 29 | 18 | [API](/api/data?q=effect&id=nausea) |
| Constipation | 12 | 7 | [API](/api/data?q=effect&id=constipation) |
| Reduced appetite | 11 | 8 | [API](/api/data?q=effect&id=reduced_appetite) |
| Vomiting | 11 | 7 | [API](/api/data?q=effect&id=vomiting) |
| Diarrhoea | 10 | 5 | [API](/api/data?q=effect&id=diarrhea) |
These five are the gastrointestinal effects most consistently reported in GLP-1 trial literature, which is exactly why they clear the bar first — trial registries and peer-reviewed papers report them routinely, in numbers, as primary or secondary adverse-event endpoints. Every figure above is live-queryable at the link beside it; check our number rather than take our word for it.
The 10 effects that don't yet clear the calibration bar (as of 21 August 2026)
| Effect | Eligible citable rate points |
|---|---|
| Headache | 7 |
| Abdominal pain | 7 |
| Acid reflux | 7 |
| Injection site reaction | 7 |
| Hair loss | 6 |
| Dizziness | 6 |
| Fatigue | 3 |
| Pancreatitis | 3 |
| Gallstones | 3 |
| Emotional blunting | 0 |
One reading note: the counts in this table are *citable rate points across both of our tracks* — clinical sources and community reports that pass the same citability checks — because that combined count is what our 10-point calibration rule is applied to. The clinical-only subset for any effect (the measure used in the first table) is usually smaller, and is what the API link beside each effect reports as `rateBase.clinical`.
This is a statement about our evidence base, not about the effects themselves. A low count here does not mean an effect is rare, mild, or doesn't happen — pancreatitis and gallstones are recognised serious risks of GLP-1 therapy regardless of how many trial abstracts state a numeric rate for them. It means that, as of today, we have not found enough independently-sourced, citable, numeric rates to run a fully calibrated statistical model for that specific effect — most commonly because the source material discusses the effect qualitatively (case reports, mechanism papers, warnings) rather than stating a comparable incidence percentage. We keep collecting daily; an effect moves from this table to the one above the day its 10th independent, citable rate point lands. Eight of these ten still have a published pooled rate from the corpus today, each labelled with a confidence grade down to "very low" for a single-source case — see the companion piece for the actual numbers. The other two, fatigue and emotional blunting, have zero eligible clinical points so far, so both the predictor and the API fall back to a labelled static literature range instead of a corpus rate. You can check any of these effects yourself via the public API (for example, pancreatitis) — it returns whichever of the two currently applies.
Why publish the gap at all
Because a database that only shows you what it's confident about is easy to mistake for a database that has answered every question. We'd rather a researcher, journalist, or patient see "3 eligible points, not enough to calibrate" for pancreatitis than assume silence means we checked and found nothing to report. The full picture — what's well-evidenced, what isn't yet, and the exact rule that draws the line — is what makes the 5 numbers above worth citing in the first place.
See the complete, current picture — all 15 effects, live counts, confidence intervals, and every source — at magistra.health/en/data-api, or read the full eligibility methodology at magistra.health/en/methodology.
See your own numbers
Our free predictor estimates your side-effect risk and weight trajectory from 1,200+ sourced data points. No signup required.
Open the predictor