Constraints and requests
Stage 1 already produced a publishable number on workstation-scale data. Getting past it needs a genomic cohort larger than any public dataset offers, and a pipeline reprocessed at a scale a workstation can't reach on its own. Neither of those is something better modeling can substitute for — they are questions of access.
We're applying to NVIDIA Inception for two things, and both follow directly from what Stage 1 already showed us. Four additional hypotheses for improving on the baseline fitness predictor were tested, and all four scored worse than the training-fold mean — so the near-term limit isn't modeling technique, it's data. Public longitudinal cohorts for clonal haematopoiesis are already exhausted, and roughly half of all CHIP mutations — frameshift, splice, and indel variants, including ASXL1 frameshifts, among the strongest signals in myeloid disease — can't be embedded as single-residue substitutions under a protein-language-model approach at all. Closing that gap means a larger cohort and, eventually, a different embedding approach. Neither is something we can do alone.
The Stage 1 pipeline already runs on NVIDIA BioNeMo's ESM-2 protein language model, cross-checked against a HuggingFace backend, on the cohorts harmonised so far. Reprocessing that same pipeline against a larger, biobank-scale cohort is a question of compute headroom to scale, not a new modeling technique.
The next scale of validation requires biobank-scale genomic cohort access, which requires an institutional partnership — not a modeling breakthrough. We don't have a confirmed partner yet. What we're asking for is the credibility and network of NVIDIA Inception to help open that conversation.
Each stage is gated on the previous one producing a number we would be willing to publish.
Lothian Birth Cohorts and SardiNIA harmonised into 591 clonal trajectories across 268 unique missense variants; signal enrichment measured at 4.4× (Lothian) and 4.1× (SardiNIA) over the noise floor; gene-level growth rates recovered blind, matching the published ordering in the literature on this same data; pipeline running on NVIDIA BioNeMo's ESM-2, cross-checked against a HuggingFace backend.
Already produced — this is the completed baseline the rest of the roadmap builds on.Reprocessing the same fitness-measurement pipeline against a genomic cohort far larger than any public longitudinal dataset offers, to test whether variant-level fitness prediction holds up at a scale that could eventually matter clinically.
Gated on securing an institutional partner.Testing whether clonal fitness prediction changes how an individual carrier is actually monitored, rather than only whether it is measurable in a research cohort.
Gated on the Stage 2 result.