Skip to main content
frontier

50% Boost in MHC‑I Antigen Presentation Predicted by AI Model

By Rachel Kim

CellType's LLMs Learn Cellular Grammar to Slash Drug‑Development Timelines

What if a model could read a cell the way it reads a sentence?

On October 17, 2025, Google Research, Google DeepMind, and Yale University jointly released a 27‑billion‑parameter foundation model called Cell2Sentence‑Scale 27B (C2S‑Scale). It is the technical spine of CellType, a Y Combinator‑backed (Winter 2026) startup building its production stack on top of it, and the basis for a pitch that the slowest stage of drug development, preclinical characterization, can be compressed by AI‑driven simulation.

C2S‑Scale was built on Google's Gemma family of open models. Where most single‑cell tools require hand‑crafted rules to identify cell types or interpret perturbations, the model "learns the hidden structure of cellular and molecular state," absorbing statistical regularities the way a text LLM absorbs syntax. A cell's gene‑expression profile becomes a sentence: a ranked sequence of genes the model reads in order and predicts forward. When a drug or genetic perturbation is introduced, the model forecasts how that sentence changes. Researchers computationally predicted and experimentally validated a novel hypothesis: that silmitasertib (a CK2 inhibitor) combined with low‑dose interferon produces a synergistic amplification of MHC‑I antigen presentation, roughly a 50% increase in the wet‑lab readout. Google CEO Sundar Pichai highlighted the result as a cover story on the work.

The grammar metaphor is not marketing. C2S‑Scale treats cells as a language because the underlying architecture was already a language model. The Gemma‑based transformer was pre‑trained on a corpus in which each cell's transcriptome is serialized into a token sequence: the higher‑expression genes first, in a rank‑ordered "sentence" of cell identity. From that representation, the model answers prompts analogous to natural‑language questions: predict the cell state after this perturbation; classify this cell type; generate a candidate gene signature. Because biological foundation models "follow clear scaling laws — just like with natural language, larger models perform better on biology," the jump from earlier million‑parameter single‑cell classifiers to a 27‑billion‑parameter foundation model is not cosmetic. It is the same curve that took GPT‑2 to GPT‑4: more data, more parameters, more emergent predictive power.

That scaling matters because the combinatorial space of drug‑gene perturbations is "experimentally unfeasible" to probe exhaustively in a lab. Every cell line, every concentration, every combination is a separate wet experiment. A model that simulates the cell‑state shift in silico collapses that combinatorial explosion into something a small team can iterate on in hours rather than months. CellType was founded by computational biologist David van Dijk, a Yale professor with 11,000‑plus citations who invented the underlying Cell2Sentence approach, and Ivan Vrkic, who co‑developed the core technology at Yale and previously led foundation‑model training elsewhere.

The architecture choice is deliberate. Earlier cell‑type annotation tools that bolted LLMs onto single‑cell data, like AnnDictionary, which uses Claude and other general LLMs to verify cell‑type mappings, hit 80–84% agreement with manual annotation on most major cell types but degraded sharply on low‑heterogeneity datasets, where "over 50% inconsistency" remained. C2S‑Scale trains natively on cellular sequences rather than adapting a general‑purpose chatbot, which is the architectural move that lets it shift from annotating cells to predicting what drugs will do to them.

From Prediction to Timeline Compression in Preclinical Work

The economic pressure on drug development is no longer abstract. AI‑driven biotech funding set a record‑breaking pace in 2025 as pharma faced flatlining pipelines and R&D budgets that bought fewer approved molecules per dollar. CellType's bet is that the bottleneck is not chemistry or clinical operations; it is the early filter, where decisions about which candidates deserve a Phase I still rest on incomplete cell biology. By turning cellular response to drug perturbations into something a 27‑billion‑parameter model can predict in silico, the company targets the single stage of drug development where timelines stretch longest.

The mechanistic argument runs through perturbation prediction itself. Single‑cell perturbation experiments have been the gold standard for decoding how a drug rewires a cell's state, but exhaustively screening combinations of compounds, doses, and cell types is "experimentally unfeasible" at the bench. CINEMA‑OT, a Nature Methods method from 2023, runs counterfactual matching across roughly 50,000 cells in about a minute, orders of magnitude faster than wet‑lab screening, and attributes treatment effects more cleanly than deep‑learning baselines. CellType inherits that lineage and scales it: a foundation model that has seen enough perturbation data to generalize the grammar of how cells respond, rather than fitting each new experiment. Early drug‑response questions that used to require a perturbation screen can now be answered computationally before a pipette touches a plate.

Recent published benchmarks make the compression concrete. XPert, a 2026 dual‑branch transformer designed for gene‑perturbation and dose–time dynamics, outperformed the next‑best baseline (TranSiGen) by 8.2% on warm‑start predictions, 15.9% on cold‑drug generalization, and 36.7% on cold‑cell generalization — the last being the regime that most closely resembles novel indications and rare cell populations. CRISP, published in Nature Computational Science in early 2026, went further: it zero‑shot predicted sorafenib's therapeutic effect in chronic myeloid leukemia from solid‑tumor training data, and the predicted mechanism (CXCR4 pathway inhibition) was corroborated by independent studies. Each of those wins translates, in operational terms, into fewer in vitro follow‑up screens, fewer dead‑end indications, and faster progression of surviving candidates into formal preclinical packages.

That framing matters because the preclinical phase is where the calendar hemorrhages. Conventional drug programs spend years in target validation, lead optimization, and IND‑enabling studies before a first human dose; any tool that collapses the front of that sequence, the hypothesis triage and indication‑selection steps, buys back months that no downstream operational reform can recover. Whether CellType's 27‑billion‑parameter model sustains that advantage across drug classes, not just oncology, is the open question the next section takes up, but on current evidence, the time compression is no longer hypothetical.

Senhwa's Bet: One Asset, Six Months, Broad Scope

Senhwa Biosciences (TPEx: 6492), a clinical‑stage oncology and rare‑disease developer, made CellType's value proposition concrete in March 2026 when it formalized a strategic memorandum of understanding built around CX‑4945, a small‑molecule CK2 inhibitor Senhwa had already advanced into the clinic. The framing of that deal matters: Senhwa did not buy a software tool. It committed to use CellType's perturbation model as a strategic repositioning engine for an existing asset, betting that AI forecasts can rewrite which indications a compound is taken into, what biomarkers gate patient selection, and which combinations enter the pipeline next.

The pilot runs only six months, but the scope is unusually broad for a pharma‑AI pilot. Under the MOU, the two companies execute on indication‑expansion strategies, biomarker discovery and validation, combination‑therapy synergy evaluation in immuno‑oncology, identification of new cancer targets, and the establishment of an "AI‑driven translational validation framework," a workflow layer most biotech pilots skip. Senhwa's public statement goes further, describing CX‑4945 as evolving "from a single‑target small molecule into a platform‑enabling asset" branded "CX‑4945 2.0." That is a corporate strategy shift, not a vendor contract: the asset's valuation logic is being rebuilt around what the model predicts it can do, not around the single mechanism that originally justified it.

The scientific hook is what made Senhwa move. During the underlying research at Yale with Google DeepMind, researchers computationally predicted and then lab‑validated a novel immune‑modulatory mechanism for CX‑4945: amplification of immune activation, enhanced tumor antigen presentation, and conversion of immunologically "cold" tumors into "hot" ones. AI‑predicted findings were subsequently validated in interdisciplinary laboratories at Yale, and CX‑4945 had originally been identified by screening more than 4,000 compounds. For Senhwa, that wet‑lab confirmation was the proof point that justified handing the asset to a model rather than treating the model as a screening filter.

"By integrating CellType's cutting‑edge AI foundation models — capable of reasoning about biology at the cellular level, and rooted in pioneering research from Yale University and Google DeepMind — into Senhwa's clinical‑stage core asset, we are not simply accelerating development; we are redefining the strategic positioning of CX‑4945." — Senhwa Chairman Benny T. Hu, March 3, 2026

Senhwa has explicitly cast itself as CellType's "founding strategic partner," and CellType has reciprocated by naming CX‑4945 the flagship case study of its platform. In return, CellType gets a real clinical asset to point at when pitching global investors, and Senhwa gets a credible AI halo for international pharma licensing discussions.

The deal also carries structural options designed to absorb exactly the kind of pipeline adjustments the model is meant to surface. The agreement "preserves flexibility for deeper collaboration structures in the future, including joint ventures, co‑development, or licensing arrangements," meaning that if the six‑month pilot identifies a high‑confidence new indication or combination, Senhwa and CellType can move quickly to fund and run it together without renegotiating from scratch. For a clinical‑stage company weighing whether to launch an additional trial arm, having a built‑in escalation path to a co‑development vehicle changes the calculus on which indications to greenlight. Whether Senhwa is a leading indicator or an outlier will become clearer as CellType signs additional pharma partners, but the structure of this first deal (single asset, six‑month pilot, broad scope, structural flexibility, reciprocal branding) is now the template rivals will be measured against.

Where the Model Stops: Wet Labs, Annotation Errors, and the FDA

C2S‑Scale 27B can forecast how a perturbation rewrites a cell's "gene sentence," but the forecast is still a forecast. The model does not run the experiment, file the IND, or sign the clinical‑trial protocol. Three layers of human work remain non‑negotiable: wet‑lab confirmation of every nominated target, regulatory validation under existing bioanalytical frameworks, and expert judgment on cell‑type annotation where the published error rates are large enough to derail a pipeline.

The wet‑lab floor is the most concrete. A 2025 review of AI‑driven virtual cell models in npj Digital Medicine put it plainly: "Mechanism‑anchored validation remains sparse; only a few virtual cell predictions have been confirmed via targeted interventions that verify model‑nominated key regulators." That sentence captures where CellType's pipeline begins. The C2S‑Scale 27B hypothesis about cancer cellular behavior, generated with Google DeepMind and covered by Pichai, was "experimentally validated in living cells" only after the model returned its prediction. Senhwa Biosciences, in its March 2026 MOU with CellType, describes the work as "AI‑driven deep data validation" intended to "expand potential indications" for CX‑4945 2.0; language that concedes the AI output is a candidate list, not a clinical asset.

Underneath every AI prediction sits a data layer with its own failure modes. Single‑cell annotation errors are the documented landmine. A 2025 Cell Biology technical guide quantifies the cascade: misidentification rates of 15–30% in cross‑tissue atlas projects; up to 50% false‑positive differential‑expression genes when annotation is 20% wrong; trajectory‑inference failure in over 40% of cases with poor annotation; and, in one in‑silico endothelial screen, a roughly 70% reduction in hit rate when the starting cell labels were wrong. The same review concludes that "incorrect annotation can lead to misinterpretation of disease biology, misidentification of therapeutic targets, and ultimately, clinical trial failure." A foundation model that ingests those labels inherits those error rates.

LLM‑based annotators do not solve the problem; they reshape it. The LICT paper in Communications Biology reports a multi‑model integration strategy that cut mismatch rates from about 22% to under 10% on PBMC data and from 11% to 8% on gastric cancer data, yet the same authors flag that "performance of LLMs diminishes when annotating less heterogeneous datasets," that "the standardized data format encoded within LLMs limits their ability to adapt to the dynamic and complex nature of biological data," and that "manual annotation benefits from expert knowledge but is inherently subjective and highly dependent on the annotator's experience." Net: automation moves some of the bias around; it does not delete it.

Then there is the cell‑culture stack the model never sees. A 2014 British Journal of Cancer guidelines paper still cited as the field reference documents "misidentification of cell lines, occult contamination with microorganisms (especially mycoplasma) and phenotypic drift due to serial transfer between laboratories" as recurring, expensive failure modes — researchers have worked for years on HeLa‑contaminated lines thinking they had KB, Int‑407, WISH, Chang liver, or Hep‑2. STR profiling is the standing authentication method (ASN‑0002 2011), with a recommended "never grow a cell line continuously for >3 months or 10 passages" before returning to frozen stock. A 27‑billion‑parameter model trained on raw counts cannot tell which of its inputs were already contaminated.

Regulation closes the loop. The FDA's draft guidance on AI to support regulatory decisions and the ICH M10 bioanalytical method validation guideline (endorsed Step 4, May 2022) require accuracy within ±15% of nominal concentration (±20% at the LLOQ) and precision under 15% CV — numbers that have to be measured in matrices, not predicted in embeddings. CellType's outputs accelerate the selection of candidates; they do not replace PK/TK runs, bioanalytical validation, or the IND review.

The practical boundary, then, is a handoff: the model ranks, the wet lab confirms, the regulator adjudicates, and a human principal investigator owns each handoff in writing. The grammar CellType is teaching its model to read is the same grammar every downstream reviewer will be asked to verify by hand — a sentence the model can write, but only a wet lab can prove.


AI Biotech Roles Hiring Now

Company Role Location Salary Range (USD/year) Roles Added (7 days)
Stripe Machine Learning Engineer South San Francisco, CA $212,000–$318,000 81
ASML Senior mixed-signal electrical engineer San Jose, CA, USA $165,375–$248,063 56

Hiring data refreshed daily from the Zero G Talent board.


Working in frontier tech? Zero G Talent tracks the openings: see every open ASML role, browse frontier tech jobs, openings at Stripe, and the people building the field.

Ready to Start Your Space Career?

Browse frontier jobs and find your next opportunity.

View frontier Jobs