01 — Structure
A protein is not finished when it is made.
After synthesis the cell decorates it. Sugars are attached and remodelled. Phosphates go on and come off in seconds. Ubiquitin marks it for destruction.
None of this is decoration. It is the control system — fast, reversible, combinatorial, and the layer at which most therapeutics act.
- GlcNAc
- Mannose
- Galactose
- Sialic acid
- Phosphate
02 — Resolution
One number per gene cannot represent this.
The same protein carries many modifiable sites, and under a single perturbation they move in opposite directions at once.
Collapse them to one abundance value and the signal cancels. The site is the unit that matters.
03 — Timescale
Sequencing is structurally blind to it.
Modifications are chemical events on the protein. They leave no trace in the transcript, so no amount of sequencing depth recovers them.
And they happen first. An RNA readout reports a consequence, later and smaller, rather than the event itself.
04
So why is every model built on RNA?
-
Scale
Sequencing produced corpora of a size modern architectures could exploit. The modelling community, reasonably, went where the data was.
-
Cost
The modification layer needs mass spectrometry and enrichment chemistry specific to each mark. Expensive per sample, and not trivially comparable between labs.
-
Result
Deep one-off studies instead of a large uniform perturbation-resolved corpus. Excellent measurement technology, almost no data in a form models can use.
05
What changes if this works.
Mechanism becomes legible
Not just that a perturbation changed the cell, but which control points moved, in what order, at what dose.
Off-target effects become visible
Consequences that never register as a change in transcript abundance stop being invisible until the clinic.
Virtual cells gain an effector layer
Models built to simulate cellular response acquire the layer where the response is executed.