The programme
An anchor disease, and a substrate built to outgrow it.
The programme is anchored on CASK-related disorder — a neurodevelopmental condition with a few hundred known patients worldwide. CASK is the first test case, not the scope. A tool that only works for one disease has failed at the job, so every instrument takes the disease as a parameter and every reference dataset is whole-genome.
Measured
Where the substrate stands
- 3,275
- papers appraised
- 2,969
- adversarial panel runs
- 50
- evidence corpora
- 144,625
- measurements extracted
- 2,629
- papers yielding measurements
- 3,112
- disease genes spanned
Each read by a three-model adversarial panel, not summarised by one.
Three frontier models extract independently, then challenge each other across rounds.
Assembled around anchor diseases, mechanisms and delivery routes.
Physical values with their units, sample size, assay and direction — not sentences about them.
A paper enters the lake only where it reports a number we can identify and compare.
The substrate is gene-general by design; a tool that only works for one disease has failed.
Derived from the research substrate on 28 August 2026 · framework 0.43.2. These are counts of our own work — papers read, panels run, measurements extracted — never a republication of any source’s content.
What it is for
Three questions the substrate is built to answer
Each is a different instrument over the same evidence base, and the third is the only one that can catch the programme being wrong.
Filter
What is worth doing
Transfer
What may be believed provisionally
Consistency
Whether it holds together
Beyond the literature
Where the programme goes past reading papers
Extraction is the foundation, not the ceiling. Several layers sit on top of it — all of them hypothesis-generating, all of them upstream of the bench.
Mechanism and cross-disease connection
Quantitative models of the disease process
Therapeutic strategy assessment
Computational molecule and binder design
Honesty
What is open, stated plainly
A programme whose claim is calibration owes you its own error bars first.
- Human review runs well behind the machine output. That is the design — arbitration trails and never blocks — but the practical consequence is that essentially everything currently surfaced is provisional and is labelled so.
- Coverage is deep in a narrow place. It is richest around the anchor disease and the neurodevelopmental genes nearest to it, and thinner as you move away. The framework reports the gap rather than filling it.
- Most measurements are not yet comparable across genes. Making a value in one paper genuinely comparable to a value in another is the hard part, and a large share of the substrate has not yet cleared that bar. Those measurements are held and counted, not discarded — and not silently included either.
- The generative layers have no proven forecasting skill. They propose; they do not predict. That is why every hypothesis carries the cheapest experiment that would falsify it.
- Nothing here has treated anyone. This is in-silico research. It becomes meaningful for a person only after laboratory and clinical scientists have tested it — and we work with researchers who do exactly that.
Working on a rare genetic disease?
The substrate is gene-general by construction. If your disease is one of the thousands it already spans, most of the machinery points at it the day you arrive.
Research aid only — not medical advice. Piper is a research and educational aid. It is not medical advice, not a diagnosis, and not a clinical determination. Its output is hypothesis-generating and may be incomplete, provisional, or wrong, including AI-generated errors. Always consult your own qualified healthcare professional before making any medical or treatment decision.