Walkthrough

Two variants, one workflow, two very different endings

The method, shown on two variants rather than described. One ends at a tablet that is already on a pharmacy shelf. The other ends at a laboratory bench, which is where most of them end. Both are public biology, neither is a result we have produced, and every step below carries its status.
01Why start with an answer we already know?

The first case is a check, not a discovery

Variants in KCNJ11 and ABCC8 cause permanent neonatal diabetes, and the established clinical answer is an oral sulfonylurea rather than lifelong insulin. About 90 percent of patients transfer off insulin successfully, and in an international cohort of 81 patients followed for a median of ten years, 75 of them, 93 percent, were still on the tablet alone. None of that is ours. It is textbook pharmacology, and it is cited.

Which is exactly why it is the case to run first. The answer is known, so it can be marked. If the workflow cannot recover the mechanism that explains it from the variant alone, the tool is not ready for the cases where nobody knows the answer.

02The case that ends in a tablet

A variant in KCNJ11

Input: gene KCNJ11, variant p.(Arg201His)

  1. S0

    Resolve and validate

    The gene symbol is checked before any inference runs, because an unrecognised symbol becomes a placeholder silently and every number after it would describe the placeholder rather than the gene. The transcript maps the coding change to a protein position, and the exact sequence that will be submitted is shown, with the substituted residue marked. Nothing is guessed on the researcher's behalf.

  2. S1

    Build the partner panel

    Kir6.2, which KCNJ11 builds, does not act alone: the channel is four Kir6.2 subunits together with four SUR1 subunits, which ABCC8 builds. That panel is assembled from public interaction databases, and eight to ten unrelated proteins are added as negative controls. A model that says yes to everything is useless, and the only way to see that is to ask it about pairs that do not interact.

  3. S2

    Wild-type map, and the first stop condition

    Every pair in the panel is predicted and set beside what is already known. Kir6.2 and SUR1 should read as a strong interaction, because they form the channel. The unrelated controls should read as weak. If the first fails, the model does not know this biology. If the second fails, the model agrees with everything. Either way the workflow stops here and says so.

  4. S3

    What the variant does

    For this variant, the change in binding energy against each partner: how much one amino-acid substitution alters each interaction.

  5. S4

    Calibration, and the second stop condition

    The same prediction is run over common, harmless variants in the same gene, to build a distribution to judge against. The question is not whether the number is large but whether it is unusual for this gene. A variant sitting inside the harmless spread is reported as no evidence of disruption, which is a real and useful answer.

  6. S5

    Molecule linkage, and the direction that is safe

    Here is the step people expect to be a prediction, and it deliberately is not one. The question asked of the corpus is whether any protein in this panel is the documented target of an approved molecule. SUR1 is the binding target of the sulfonylureas. That is established pharmacology, retrieved from public sources as fact with its date attached, not generated by a model.

  7. S6

    Filters that turn a molecule into a candidate

    Approved and available where the team works. Paediatric experience, from the label and the dosing literature. The safety record, retrieved as observed clinical fact. Whether the molecule reaches the tissue that matters, which is irrelevant here and central to the second case. Glibenclamide clears all of it: a decades-old, inexpensive generic.

  8. S7

    The report, and where Biofacta stops

    Out comes the variant, the disrupted interaction, the affected complex, the documented pharmacology of the implicated partner, the safety and availability record, and, attached to every model-derived number, its caveat and its reference figures as received, travelling into the export rather than staying on the screen.

What Biofacta does not say is give this child glibenclamide. It says this variant is predicted to disrupt this interaction in this channel, this partner is the documented target of this approved class, here is the evidence, and here is what each piece of it is worth. A clinician decides.

03The case that ends in a laboratory

A variant of uncertain significance in COMMD9

Same seven steps. A different, and far more typical, ending.

The panel is not a guess

Commander is a sixteen-subunit assembly: COMMD1 to COMMD10 in a ring, scaffolded by CCDC22 and CCDC93, plus DENND10 and the Retriever trimer. Extended with the associated sorting machinery, that is about twenty-five proteins, and every pair is a defined question rather than an open one.

And then the workflow returns nothing

No approved molecule documents this kind of endosomal recycling machinery as its target. The corpus is silent, and Biofacta says so plainly. It does not fill the gap with a weak number, and it does not present an empty result as a system failure.

What the team gets instead, which is not nothing

  • A mechanism for a variant that was previously uninterpretable, which moves it toward a classification and may end a diagnostic odyssey.
  • The specific interaction predicted to break, which is a bench experiment measured in days rather than months.
  • Structural context: where on the interface it acts, which is what a reviewer accepts as explanation rather than as prediction.
  • A cohort, because other patients with variants in the same complex become findable.
  • Provenance a publication can carry: every number traceable to its inputs, its model version and its caveats.

Most rare-disease variants do not end at a tablet. They end at a mechanism a laboratory can test.

04What does a study like this cost to run?

Minutes, and the expensive step is the one we do not perform

For the whole twenty-five protein study, counted in model requests:

StepRequests
Validate about 25 gene symbols25
All-pairs interaction map across 25 proteins300
Negative controls200
Variant effect, about 10 variants against 15 partners150
Harmless-variant calibration600
Determinism repeats50
Total~1,325
A broad drug-repurposing sweep~42,000 requests · ~3 hours
This study, cheapest checks first~1,325 requests · ~6 minutes
Our own planning figures for the twenty-five protein study above. The sweep produces a filter rather than a ranking, which is why the comparison is about cost and not about quality.

About six minutes. The broad drug-repurposing sweep this replaces is roughly 42,000 requests and three hours, and it produces a filter rather than a ranking. The step everyone assumes is the expensive one is the step Biofacta does not perform.

05What of this runs today?

The honest status of every step

Marked plainly, because the alternative is that you find out in a meeting.

  1. S0

    Resolve and validate

    Runs now
  2. S1

    Partner panel from public interaction data

    Runs now
  3. S2

    Wild-type interaction map

    The endpoint exists; what serves it and how good it is are not yet independently checked.

    Exists, quality unverified
  4. S3

    Variant effect on each interaction

    The strongest task in the underlying published work has no public checkpoint. We are training it.

    Not yet built
  5. S4

    Calibration against harmless variants

    Depends on the step above.

    Not yet built
  6. S5

    Molecule linkage from the corpus

    Runs now
  7. S6

    Filters and regional availability

    The tissue-reach model is published but not yet loaded; availability data needs sourcing per region.

    Exists, quality unverified
  8. S7

    Report and provenance

    Runs now

Read the pattern rather than the rows. What works today is the data and the provenance. The step that makes the workflow worth having is the variant-effect step, and it does not exist yet. That is the whole project in one table, and it is the reason we are building that step first rather than last.

06Where does a person decide?

Biofacta narrows. People decide.

Every gate below is a stop condition rather than a warning. The tool is designed to halt and explain rather than to produce something for every input.

PointDecisionWho
InputIs this the right transcript and variant?Clinical geneticist
PanelIs this partner panel biologically appropriate?Scientific lead
First gateDid the model recover known biology? Stop if not.Scientist
Second gateDoes the method discriminate on this gene? Stop if not.Scientist
LinkageIs the documented pharmacology relevant to this mechanism?Clinician or pharmacologist
ReportIs there a case here at all?Multidisciplinary team
07What can this workflow never do?

Stated every time, not once in the small print

CannotWhy
Predict whether a patient tolerates a medicineNo pharmacokinetics and no adverse-event capability. Safety comes from the empirical record, never from a prediction.
Predict clinical benefitBinding is not efficacy. Nothing here predicts whether a patient improves.
Rank medicines against a targetThe relevant axis is flat. Refused by design, with the reason shown where you look for it.
Say anything about loss-of-function variantsNo protein, no interaction, no structure. Frameshift, nonsense, splice and copy-number findings are out of scope, and we say so rather than quietly excluding them.
Account for cell type, tissue or developmental timingNo expression or localisation capability.
Replace a diagnostic testResearch use only. Not a medical device, and no clinical claim.

Bring us the case you cannot explain

The second case above is the one we expect to see most. If you have a protein you already care about and a variant that has not resolved, tell us about it.

Bring us a variant