Need help right now? Lifeline 13 11 14 Suicide Call Back 1300 659 467 Poisons 13 11 26 Emergency 000
The evidence

How strong is this, actually?

Every number on this page opens its source, the quote it came from, and our own tier for it. Where we could not resolve a claim to a primary paper, the source line says so instead of dressing it as one.

Tiers

Six tiers, and one refusal

The point is not that a link exists. Most of them do, somewhere, in some study. The point is how much weight it can hold. Tap any tier for what it takes to earn it.

A correction slip pasted flat over the lines of a printed page it corrects, its bottom edge lifting. A correction slip pasted flat over the lines of a printed page it corrects, its bottom edge lifting.
The Withdrawn tier is this object: a correction physically attached to the thing it corrects, covering the line it retracts. Most sources never paste the slip in.

Below all six sits Insufficient, a refusal rather than a tier. If nobody has properly checked, that is what you are told.

One decision, in full

We stopped using the word "rare"

Product information sorts side effects into bands: very common, common, uncommon, rare, very rare. We used to show the band and the number, because that felt obviously more informative. It was the worst of the three options that have been tested.

13in 100 may get this
87in 100 do not

Under 1% the number stops helping, so it is drawn instead. Thirteen in a hundred, with the eighty-seven that are not, at the same weight.

What the trials found

Adding the word to the number pushes the estimate of "any side effect" from 18.7% to 31.1%. In one study only 7 of 180 people placed "common" and "rare" inside the ranges those words officially carry.

The sharpest version: shown "rare" against a stated 0.04% risk of pancreatitis, patients estimated 18.0%. The word overrides the arithmetic beside it by a factor of hundreds.

Positive framing cuts attribution from 54.5% to 39.2% without reducing how many symptoms people have, in participants who were not patients. So both halves render at the same size.

The obvious objection is that telling people at all is the harm. The evidence runs the other way: a leaflet improved adherence (OR 3.0), and belief that a medicine is necessary drives taking it while worry drives stopping. Withholding is not the safe option. Framing is the only lever there is.

The research

Ten runs, and the parts that proved us wrong

Before building most of this we ran ten deep literature and regulatory reviews: Australian medical-device law, prescription-medicine advertising, health-record access, data licensing, privacy, peer-forum duty of care, daily self-report methodology, nocebo and risk framing, crisis-surface design, and the licensing of medication-burden scales. Not every backend returned a usable report; the ones that landed are published unedited, including the parts that contradict us.

0research runs commissioned
0evidence tiers, plus a refusal
0daily items, none of them scored
0findings that resolve to an action
The compressed edge of a tall stack of printed paper, with plain tabs protruding at intervals. The compressed edge of a tall stack of printed paper, with plain tabs protruding at intervals.
Ten runs, printed. The tabs mark the pages that contradict us, which is the only reason publishing them unedited means anything.
We assumedThe research found
The crisis screen is the safe partIt is the part most likely to make the whole app a regulated medical device, because routing you on your symptoms is triage, and a single regulated feature pulls in the entire application
The symptom lookup is the risky partIt probably survives if it stays a text search of published information and never ranks by likelihood
A prescription QR can be read offlineIt cannot. It carries a Delivery Service Prescription Identifier, not the medicine
Screening posts before publication is the safe moderation choiceIt decisively increases the duty of care we would owe. It is also the best-evidenced safety mechanism. A real trade-off, not a mistake
Four assumptions were backwardsFour, from the first five runs. The later runs inverted roughly half of what we brought them, which is the actual number
Base rates reassure anxious readersThey fail exactly where they are most needed: salience reduced base-rate neglect in healthy participants and did nothing for people with elevated health anxiety

Our own check-in lost half its questions.

Water intake went because the within-person evidence is not there. Hours slept went for a different reason: the evidence does exist, at roughly 0.04–0.05 SD on next-day affect, but sleep quality beats it and we were asking both. Cut for redundancy is a different finding from cut for absence, and we had been reporting the first as the second.

Sources

How solid is any of this?

Every marker on this site opens the quote, the source and the tier we give it. What that mechanism cannot do is audit itself, so here is the state of it plainly.

What has been checked

Each claim carries its own source, taken from the run reports and their registries. Where two of our own reports disagreed we followed the more cautious reading and said so in the write-ups.

What has not

These reports were produced by AI deep-research systems and have not been through an independent claim-verification pass. Treat the figures as well-sourced and not yet audited. Two markers on this site say plainly that we could not resolve them to a primary paper, and where our own release gate refuses to print a figure, because nobody here has read the paper it would come from, this site does not print it either.

Every report is public. If you think one of them is wrong, we would genuinely like to know which part.