64% of Our GLP-1 Rate Rows Come From Six Trials That Count Some Participants More Than Once. The Rates Barely Move.
We ran an arithmetic test on our own evidence base: sum the at-risk denominators of the trial arms we store, and if they exceed the trial's own enrolment, some participants are inside more than one of our rows. Six of 27 registry studies fail it, and 726 of our 1,134 registry rate rows come from those six. Removing all six moves every published pooled rate by at most 0.5 percentage points — and moves three source-diversity grades down a step. Here is the full before-and-after table.
Every rate we publish carries an n beside it: "23% pooled, from 135 stated rates across 29 distinct studies". We put that n there on purpose, in August 2026, because a percentage with no denominator is a number wearing a lab coat.
This article is about a defect in that n. It is not about the percentages — those turn out to be fine, and we measured that rather than hoping it.
The test
A rate row from a trial registry is one ARM of one trial: "in arm X (n at risk), term T occurred in k participants". ClinicalTrials.gov posts each event group separately, and for a crossover or dose-escalation study those event groups are the same people, several times over — period 1, period 2, dose level 3, fed state, fasted state. Our store keeps each as its own row, so one cohort of 52 volunteers can supply eighteen of our "stated rates".
The test for this needs no judgement and cannot false-positive:
> For one study and one effect, sum the at-risk denominators of the arms we store. If that sum exceeds the trial's own posted enrolment, some participants sit inside more than one of our rows.
A study cannot have more people in disjoint arms than it enrolled. All six studies below are completed and post an ACTUAL enrolment count, so the denominator of the test is a final number and not a recruitment target.
The six studies
Measured against our production store on 21 September 2026:
| Study | Arms stored (worst effect) | Sum of arm n | Posted enrolment | Ratio | Live rate rows |
|---|---|---|---|---|---|
| NCT05841238 | 18 | 600 | 52 | 11.5x | 144 |
| NCT06440980 | 38 | 3,772 | 533 | 7.1x | 342 |
| NCT03929744 | 21 | 291 | 133 | 2.2x | 168 |
| NCT06370728 | 2 | 58 | 30 | 1.9x | 16 |
| NCT05891496 | 2 | 33 | 23 | 1.4x | 16 |
| NCT05086445 | 5 | 68 | 62 | 1.1x | 40 |
Six of the 27 registry studies with live rated rows fail the test. Between them they supply 726 of our 1,134 live registry rate rows (64.0%) — and because almost every rated row we now hold is a registry arm row, that is 726 of the 1,142 eligible stated rates we publish site-wide, 63.6%. NCT06440980 alone supplies 342 of them, the most of the six; NCT05841238 has the worst ratio, 600 arm-participants recorded across 52 actual ones.
None of these rows is wrong. Each is a true rate, correctly attributed, quoted verbatim from the registry's own table. The defect is in what the COUNT of them implies.
What it costs: every effect, with and without
The obvious next question is what our published numbers would look like if all six studies were removed. We did not estimate that; we re-ran the same pooling function the site publishes from with those studies' rows excluded.
| Effect | Published | Without the six | Move |
|---|---|---|---|
| Nausea | 23% — 135 rates / 29 studies (very high) | 23.3% — 49 / 23 (high) | +0.3pp |
| Diarrhoea | 17.2% — 128 / 24 (high) | 17.4% — 46 / 20 (high) | +0.2pp |
| Constipation | 11.8% — 128 / 24 (high) | 12.0% — 42 / 18 (high) | +0.2pp |
| Headache | 10.7% — 120 / 20 (high) | 11.1% — 34 / 14 (high) | +0.4pp |
| Vomiting | 10.3% — 133 / 27 (very high) | 10.5% — 47 / 21 (high) | +0.2pp |
| Reduced appetite | 9.3% — 125 / 24 (high) | 9.4% — 39 / 18 (high) | +0.1pp |
| Abdominal pain | 7.2% — 116 / 17 (high) | 7.7% — 32 / 12 (high) | +0.5pp |
| Fatigue | 6.9% — 65 / 12 (high) | 7.4% — 23 / 9 (moderate) | +0.5pp |
| Dizziness | 6.5% — 97 / 17 (high) | 6.6% — 29 / 12 (high) | +0.1pp |
| Acid reflux | 5.0% — 52 / 14 (high) | 5.0% — 32 / 12 (high) | 0 |
| Hair loss | 4.5% — 8 / 3 (low) | unchanged | 0 |
| Injection-site reaction | 4.5% — 4 / 2 (low) | unchanged | 0 |
| Gallstones | 1.7% — 20 / 7 (moderate) | unchanged | 0 |
| Pancreatitis | 0.1% — 11 / 3 (low) | unchanged | 0 |
| Emotional blunting | no corpus estimate | no corpus estimate | — |
Nine of the fourteen pooled estimates move at all, and the largest move in the table is 0.5 percentage points. Every one of the nine moves in the same direction — upward — which is what you would expect if the repeated-participant studies are mostly early-phase trials reporting lower per-arm rates than the large outcome trials.
Why the percentages survive a defect this big
Because the pooling never treated those rows as independent in the first place. Our pooled estimate collapses each source to one entry — the unweighted mean of the arms it posts, carried at the size of its largest single arm — and then weights the entries by that stated sample size, multiplied by a discount for how confident our extraction was in reading each figure. So NCT06440980, which posts 38 nausea arms over 533 people, enters the weighted mean once, at n=215 — the size of its largest single arm, not the trial's enrolment and not the 3,772 its arms sum to — exactly as it would if the registry had posted that one arm and nothing else.
For thirteen of the fifteen effects, one entry per source is also exactly one entry per study. The two exceptions are in the table above: pancreatitis and gallstones, where a trial can post two MedDRA preferred terms for the same effect and so enters the mean twice. Pancreatitis runs to 5 entries across 3 studies — SURMOUNT-1 and SELECT each post "Pancreatitis" and "Pancreatitis acute" — and gallstones to 10 across 7, where SURMOUNT-1, SELECT and STEP 1 each post "Cholecystitis" and "Cholelithiasis". None of the six studies this article is about contributes a row to either effect.
So the arms inflate the row count and not the estimate. That is a real property of the method and not an accident, but until today nobody had measured how far the two diverge, and a reader had no way to know.
What does move: the n, and three grades
Two things change materially.
The n itself. Nausea's base is 135 stated rates across 29 studies; without the six it is 49 across 23. Vomiting goes 133/27 to 47/21. Abdominal pain, dizziness and headache each lose about seven rows in ten (116 to 32, 97 to 29, 120 to 34). A reader comparing "135 rates" for nausea against "8 rates" for hair loss is comparing two quantities that do not mean the same thing at the same scale.
Source diversity. Our sourceDiversity label is derived purely from the distinct-study count, so removing six studies moves three effects down a grade: nausea very high to high, vomiting very high to high, fatigue high to moderate. Those are the labels a reader skims, and they are the ones this defect flatters.
What we changed today, and what we have not
From today, every /api/data response carries an armOverlapNote field stating the test, the six studies, the 64% figure and this before-and-after summary; the CC BY aggregate table carries the same text as countCaveat, so a redistributor gets the caveat with the columns; and both measurement scripts are published in our methodology mirror.
What we have not done is change any published number, because it is genuinely not obvious which answer is right. Excluding the six discards real, correctly-reported adverse-event data from six real trials in order to fix a count. Keeping them and disclosing the overlap keeps the data and asks the reader to hold a caveat. A third option — collapsing a study's overlapping periods into one row before counting — is the most defensible and the most work, and it needs a rule for telling a genuine parallel arm from a repeated period, which the registry does not always make explicit.
We have not picked one. The honest position today is that our percentages are robust to this and our counts are not, both are now stated on the same endpoint, and the decision is open.
The thing to take away
If you are reading anyone's real-world-evidence dataset, "N rates from M sources" is two claims and the second one is the load-bearing one. Ask what a "rate" is a rate of — a patient cohort, or a row in someone's store. For ours, until today, the answer was a row, and we had not said so.
Reproduce it
Both steps are committed and published: measure-arm-overlap.mjs runs the enrolment test, and measure-overlap-exclusion-impact.mjs re-runs the production pooling function with the flagged studies removed and prints the table above. They read our own store, so they are the audit trail rather than something you can run yourself — but the test needs nothing from us. Every arm and every posted enrolment above is in the ClinicalTrials.gov record linked beside it, and the rate rows we hold for each study are listed under sources on our per-effect endpoint.
All figures read from production on 21 September 2026.
See your own numbers
Our free predictor estimates your side-effect risk and weight trajectory, with the stated rates and distinct sources shown behind every figure. No signup required.
Open the predictorWorking from the data itself? The dated snapshot behind these figures is available as a one-off purchase, alongside the free public API: Data & API.