How it works

The reasoning is the product.

A single model summarising a paper gives you one fluent opinion and no way to tell whether it is right. Piper spends its effort somewhere else: on producing a disagreement you can inspect, a measurement you can compare, and an explicit record of what the evidence will not support.

Step one

A panel, not a summariser

Three frontier models from three different vendors read the same paper and answer the same structured set of methodological questions — roughly forty of them, the ones a careful reviewer would ask.

Independent extraction

Each model answers on its own, with no sight of the others. Independence first — otherwise the second and third opinions are just echoes of the first.

Mutual evaluation

Each model then sees what the others said and is asked where they are wrong. Most disagreements resolve here, and the resolution is recorded with its reasoning.

Persistent disagreement

What survives that round is a genuine split. It is pushed once more, and then it stops being a machine problem: it becomes a question for a person.

Convergence — and suspicion of it

Agreement is computed, but fast unanimous agreement on a fact the models could have memorised is flagged as an anti-signal. Three models trained on overlapping text are perfectly capable of being confidently wrong together.

Why this is expensive on purpose. The panel costs several dollars per paper because the argument is what we are buying. The extracted value is close to a by-product of it — and when a reviewer later asks why does it say that, the whole exchange is still there.

Step two

Measurements with an identity precise enough to compare

Appraising a paper tells you how much to trust it. It does not tell you what it measured. Those are different jobs, and the second one is where most evidence tooling quietly goes wrong.

A number is not a measurement

A value is captured with its units, its sample size, its dispersion, the assay that produced it, the biological system it came from, and the direction it moved against a named comparator. Without all of that it cannot be compared to anything.

Identity is built from every axis that makes two measurements different

The property, the target, the qualifiers, the dimensional class — and the organism. A 300-gram human brain and a half-gram mouse brain are not the same measurement, and a system that pools them will produce a confident answer that is simply false.

Every rung refuses rather than reaches

Where the identity cannot be resolved, the measurement is set aside and counted, not guessed at. A wrong split costs a comparison and can be repaired. A wrong merge manufactures a conclusion about a rare disease and cannot.

A missing number is a recorded state, not a zero

When a paper does not report a value, that is stored as its own fact with a reason. An absent measurement and a measurement of zero mean opposite things, and collapsing them is how an evidence base starts lying.

Step three

Carrying what is known onto what is not

This is the whole reason the substrate is gene-general. An ultra-rare disease will never have its own literature — but the genes and mechanisms around it may.

For a measurement the anchor disease has no data on, the framework asks which better-studied genes do — and then asks the much harder question of whether their result may be read across. That is only allowed on an explicit warrant: the two genes must have moved the same way on the same measurement, in independent papers. Agreement that comes from a single experiment measuring both genes together is reported separately, because a shared protocol and a shared cohort make that agreement close to guaranteed by design.

Where the warrant does not hold, the honest output is that nothing is known — not a weaker version of agreement. Lowering the bar would restore the yield by deleting the guard that made the yield mean something.

Every one of these is a prediction, never evidence. None of it is written back into the substrate as a measurement of the recipient gene, and each one carries the experiment that would kill it: measure the thing in the anchor disease’s own system, and a result in the opposite direction ends the hypothesis.

Step four

Humans trail; they do not block

The distinction that governs everything downstream, and the one most likely to be misread.

Machine output is presumptive. It is surfaced, it is usable, and it is labelled — but it is never promoted to confirmed by any automatic rule, however strong the agreement. The only path to confirmed is an explicit human ruling, recorded with its reasoning and its author.

That review runs behind the analysis rather than in front of it, deliberately. A framework that stopped every time a person had not yet caught up would produce nothing at all in a field this thin. The consequence is one we would rather state than let you discover: the great majority of what you will see is provisional, and it is marked that way everywhere it appears.

Reviewers are chosen cross-domain on purpose — an epidemiologist and a biologist from an adjacent field will catch different failures than two specialists in the same subfield, who tend to share blind spots.

The guards

Structural, not aspirational

Each of these is enforced by code or by the database and fails the build when it is violated. They exist because each one was, at some point, a silent bug.

Conservation with attribution

At every step the arithmetic must close: input equals output plus everything excluded, and each exclusion carries a reason from a closed list that must be true of it. A step that cannot account for itself throws rather than continuing.

No silent ceilings

Any limit that drops or truncates data must announce itself. The failure this prevents is the one humans are worst at noticing — not an error, but an absence.

Identifiers are resolved, never recalled

Ontology and gene identifiers are looked up against pinned indexes. Where a label cannot be resolved it is reported unresolved and counted, never replaced with the nearest plausible code.

No default may be the most confident option

Where a value has to fall back, it falls back to the least-committal member of its vocabulary — and is recorded as defaulted, so a consumer can exclude it and an auditor can count it.

The framework cannot calibrate on itself

Its own outputs are barred from feeding its own corrective changes — enforced at the import boundary, so the path cannot be created by accident.

Licensed sources inform; they are never republished

Third-party licensed data is used in the backend only. What the service returns is our own derived conclusion, never the underlying rows — including on this page, where the figures are counts of our own work.

See it against the work

The programme page sets out what has actually been produced, and what remains open.

Research aid only — not medical advice. Piper is a research and educational aid. It is not medical advice, not a diagnosis, and not a clinical determination. Its output is hypothesis-generating and may be incomplete, provisional, or wrong, including AI-generated errors. Always consult your own qualified healthcare professional before making any medical or treatment decision.

How it works — Piper by Focena