Protein-modification intelligence

Cells are modelled from their RNA. The work is done by their proteins.

We build models and agent systems for the protein layer of the cell — the proteins themselves, and the chemical modifications that switch them on and off.

01 — The gap

A transcript is a plan. It is not the action.

Nearly every virtual-cell and perturbation model is trained on transcriptomes, because that is the measurement that scaled first. It is a useful proxy. It is still a proxy.

The work is done by proteins, governed by the chemical marks placed on them after they are made. That layer is where signalling happens, where most drugs act, and where sequencing is structurally blind.

02 — What is missing

Schematic. Response after a perturbation, by layer

The fast layer moves first, and moves most.

Signalling reorganises in seconds to minutes. Transcription responds later, smaller, and downstream — so an RNA readout reports a consequence rather than the event.

A model built only on RNA is not missing a modality. It is missing the layer where the causal action occurs.

03 — Position

Modification state is treated as an annotation. It should be something a model predicts.

In today’s structure and sequence models a modification is a lookup: a code attached to a residue index, supplied as input, never estimated. Nothing predicts how that layer moves when the cell is perturbed.

The reason is not architectural taste. The training data does not exist at the scale these models need. Data problem first, model problem second — that ordering determines how we build.

04 — What we work on

/ 01

Representation

Modification state as a modelled variable with its own objective, not metadata carried alongside a sequence.

/ 02

Agents in the loop

Systems that propose a perturbation, run the analysis, read the result, and decide what to measure next.

/ 03

The data underneath

Generated by us, because what these models need does not yet exist in a form they can learn from.

05 — The core

The technical detail is not public yet.

We would rather show the work than describe it. When there is a result worth reading, it will be here.

● Coming soon

Model, benchmark, data engine

06 — Contact

If you work on virtual-cell models, perturbation modelling, or proteomics at scale, we’d like to hear from you.

[email protected]