How it works
The reasoning is the product.
A single model summarising a paper gives you one fluent opinion and no way to tell whether it is right. Piper spends its effort somewhere else: on producing a disagreement you can inspect, a measurement you can compare, and an explicit record of what the evidence will not support.
Step one
A panel, not a summariser
Three frontier models from three different vendors read the same paper and answer the same structured set of methodological questions — roughly forty of them, the ones a careful reviewer would ask.
Independent extraction
Mutual evaluation
Persistent disagreement
Convergence — and suspicion of it
Why this is expensive on purpose. The panel costs several dollars per paper because the argument is what we are buying. The extracted value is close to a by-product of it — and when a reviewer later asks why does it say that, the whole exchange is still there.
Step two
Measurements with an identity precise enough to compare
Appraising a paper tells you how much to trust it. It does not tell you what it measured. Those are different jobs, and the second one is where most evidence tooling quietly goes wrong.
A number is not a measurement
Identity is built from every axis that makes two measurements different
Every rung refuses rather than reaches
A missing number is a recorded state, not a zero
Step three
Carrying what is known onto what is not
This is the whole reason the substrate is gene-general. An ultra-rare disease will never have its own literature — but the genes and mechanisms around it may.
For a measurement the anchor disease has no data on, the framework asks which better-studied genes do — and then asks the much harder question of whether their result may be read across. That is only allowed on an explicit warrant: the two genes must have moved the same way on the same measurement, in independent papers. Agreement that comes from a single experiment measuring both genes together is reported separately, because a shared protocol and a shared cohort make that agreement close to guaranteed by design.
Where the warrant does not hold, the honest output is that nothing is known — not a weaker version of agreement. Lowering the bar would restore the yield by deleting the guard that made the yield mean something.
Every one of these is a prediction, never evidence. None of it is written back into the substrate as a measurement of the recipient gene, and each one carries the experiment that would kill it: measure the thing in the anchor disease’s own system, and a result in the opposite direction ends the hypothesis.
Step four
Humans trail; they do not block
The distinction that governs everything downstream, and the one most likely to be misread.
Machine output is presumptive. It is surfaced, it is usable, and it is labelled — but it is never promoted to confirmed by any automatic rule, however strong the agreement. The only path to confirmed is an explicit human ruling, recorded with its reasoning and its author.
That review runs behind the analysis rather than in front of it, deliberately. A framework that stopped every time a person had not yet caught up would produce nothing at all in a field this thin. The consequence is one we would rather state than let you discover: the great majority of what you will see is provisional, and it is marked that way everywhere it appears.
Reviewers are chosen cross-domain on purpose — an epidemiologist and a biologist from an adjacent field will catch different failures than two specialists in the same subfield, who tend to share blind spots.
The guards
Structural, not aspirational
Each of these is enforced by code or by the database and fails the build when it is violated. They exist because each one was, at some point, a silent bug.
Conservation with attribution
No silent ceilings
Identifiers are resolved, never recalled
No default may be the most confident option
The framework cannot calibrate on itself
Licensed sources inform; they are never republished
See it against the work
The programme page sets out what has actually been produced, and what remains open.
Research aid only — not medical advice. Piper is a research and educational aid. It is not medical advice, not a diagnosis, and not a clinical determination. Its output is hypothesis-generating and may be incomplete, provisional, or wrong, including AI-generated errors. Always consult your own qualified healthcare professional before making any medical or treatment decision.