Walkthrough
Two variants, one workflow, two very different endings
The first case is a check, not a discovery
Variants in KCNJ11 and ABCC8 cause permanent neonatal diabetes, and the established clinical answer is an oral sulfonylurea rather than lifelong insulin. About 90 percent of patients transfer off insulin successfully, and in an international cohort of 81 patients followed for a median of ten years, 75 of them, 93 percent, were still on the tablet alone. None of that is ours. It is textbook pharmacology, and it is cited.
Which is exactly why it is the case to run first. The answer is known, so it can be marked. If the workflow cannot recover the mechanism that explains it from the variant alone, the tool is not ready for the cases where nobody knows the answer.
A variant in KCNJ11
Input: gene KCNJ11, variant p.(Arg201His)
- S0
Resolve and validate
The gene symbol is checked before any inference runs, because an unrecognised symbol becomes a placeholder silently and every number after it would describe the placeholder rather than the gene. The transcript maps the coding change to a protein position, and the exact sequence that will be submitted is shown, with the substituted residue marked. Nothing is guessed on the researcher's behalf.
- S1
Build the partner panel
Kir6.2, which KCNJ11 builds, does not act alone: the channel is four Kir6.2 subunits together with four SUR1 subunits, which ABCC8 builds. That panel is assembled from public interaction databases, and eight to ten unrelated proteins are added as negative controls. A model that says yes to everything is useless, and the only way to see that is to ask it about pairs that do not interact.
- S2
Wild-type map, and the first stop condition
Every pair in the panel is predicted and set beside what is already known. Kir6.2 and SUR1 should read as a strong interaction, because they form the channel. The unrelated controls should read as weak. If the first fails, the model does not know this biology. If the second fails, the model agrees with everything. Either way the workflow stops here and says so.
- S3
What the variant does
For this variant, the change in binding energy against each partner: how much one amino-acid substitution alters each interaction.
- S4
Calibration, and the second stop condition
The same prediction is run over common, harmless variants in the same gene, to build a distribution to judge against. The question is not whether the number is large but whether it is unusual for this gene. A variant sitting inside the harmless spread is reported as no evidence of disruption, which is a real and useful answer.
- S5
Molecule linkage, and the direction that is safe
Here is the step people expect to be a prediction, and it deliberately is not one. The question asked of the corpus is whether any protein in this panel is the documented target of an approved molecule. SUR1 is the binding target of the sulfonylureas. That is established pharmacology, retrieved from public sources as fact with its date attached, not generated by a model.
- S6
Filters that turn a molecule into a candidate
Approved and available where the team works. Paediatric experience, from the label and the dosing literature. The safety record, retrieved as observed clinical fact. Whether the molecule reaches the tissue that matters, which is irrelevant here and central to the second case. Glibenclamide clears all of it: a decades-old, inexpensive generic.
- S7
The report, and where Biofacta stops
Out comes the variant, the disrupted interaction, the affected complex, the documented pharmacology of the implicated partner, the safety and availability record, and, attached to every model-derived number, its caveat and its reference figures as received, travelling into the export rather than staying on the screen.
What Biofacta does not say is give this child glibenclamide. It says this variant is predicted to disrupt this interaction in this channel, this partner is the documented target of this approved class, here is the evidence, and here is what each piece of it is worth. A clinician decides.
A variant of uncertain significance in COMMD9
Same seven steps. A different, and far more typical, ending.
The panel is not a guess
Commander is a sixteen-subunit assembly: COMMD1 to COMMD10 in a ring, scaffolded by CCDC22 and CCDC93, plus DENND10 and the Retriever trimer. Extended with the associated sorting machinery, that is about twenty-five proteins, and every pair is a defined question rather than an open one.
And then the workflow returns nothing
No approved molecule documents this kind of endosomal recycling machinery as its target. The corpus is silent, and Biofacta says so plainly. It does not fill the gap with a weak number, and it does not present an empty result as a system failure.
What the team gets instead, which is not nothing
- A mechanism for a variant that was previously uninterpretable, which moves it toward a classification and may end a diagnostic odyssey.
- The specific interaction predicted to break, which is a bench experiment measured in days rather than months.
- Structural context: where on the interface it acts, which is what a reviewer accepts as explanation rather than as prediction.
- A cohort, because other patients with variants in the same complex become findable.
- Provenance a publication can carry: every number traceable to its inputs, its model version and its caveats.
Most rare-disease variants do not end at a tablet. They end at a mechanism a laboratory can test.
Minutes, and the expensive step is the one we do not perform
For the whole twenty-five protein study, counted in model requests:
| Step | Requests |
|---|---|
| Validate about 25 gene symbols | 25 |
| All-pairs interaction map across 25 proteins | 300 |
| Negative controls | 200 |
| Variant effect, about 10 variants against 15 partners | 150 |
| Harmless-variant calibration | 600 |
| Determinism repeats | 50 |
| Total | ~1,325 |
About six minutes. The broad drug-repurposing sweep this replaces is roughly 42,000 requests and three hours, and it produces a filter rather than a ranking. The step everyone assumes is the expensive one is the step Biofacta does not perform.
The honest status of every step
Marked plainly, because the alternative is that you find out in a meeting.
- S0Runs now
Resolve and validate
- S1Runs now
Partner panel from public interaction data
- S2Exists, quality unverified
Wild-type interaction map
The endpoint exists; what serves it and how good it is are not yet independently checked.
- S3Not yet built
Variant effect on each interaction
The strongest task in the underlying published work has no public checkpoint. We are training it.
- S4Not yet built
Calibration against harmless variants
Depends on the step above.
- S5Runs now
Molecule linkage from the corpus
- S6Exists, quality unverified
Filters and regional availability
The tissue-reach model is published but not yet loaded; availability data needs sourcing per region.
- S7Runs now
Report and provenance
Read the pattern rather than the rows. What works today is the data and the provenance. The step that makes the workflow worth having is the variant-effect step, and it does not exist yet. That is the whole project in one table, and it is the reason we are building that step first rather than last.
Biofacta narrows. People decide.
Every gate below is a stop condition rather than a warning. The tool is designed to halt and explain rather than to produce something for every input.
| Point | Decision | Who |
|---|---|---|
| Input | Is this the right transcript and variant? | Clinical geneticist |
| Panel | Is this partner panel biologically appropriate? | Scientific lead |
| First gate | Did the model recover known biology? Stop if not. | Scientist |
| Second gate | Does the method discriminate on this gene? Stop if not. | Scientist |
| Linkage | Is the documented pharmacology relevant to this mechanism? | Clinician or pharmacologist |
| Report | Is there a case here at all? | Multidisciplinary team |
Stated every time, not once in the small print
| Cannot | Why |
|---|---|
| Predict whether a patient tolerates a medicine | No pharmacokinetics and no adverse-event capability. Safety comes from the empirical record, never from a prediction. |
| Predict clinical benefit | Binding is not efficacy. Nothing here predicts whether a patient improves. |
| Rank medicines against a target | The relevant axis is flat. Refused by design, with the reason shown where you look for it. |
| Say anything about loss-of-function variants | No protein, no interaction, no structure. Frameshift, nonsense, splice and copy-number findings are out of scope, and we say so rather than quietly excluding them. |
| Account for cell type, tissue or developmental timing | No expression or localisation capability. |
| Replace a diagnostic test | Research use only. Not a medical device, and no clinical claim. |
Bring us the case you cannot explain
The second case above is the one we expect to see most. If you have a protein you already care about and a variant that has not resolved, tell us about it.
