What Incyte Bought for $120 Million
Pharmaceutical companies have spent years renting AI platforms. Incyte just paid to help build the next one.
On May 20, Incyte and Genesis Molecular AI announced an expanded collaboration that restructures the pharma-AI relationship around a single premise: the drugmaker's proprietary experimental data becomes training fuel for the AI company's foundation models. The agreement commits $120 million upfront, according to Incyte's investor press release: $80 million cash, Genesis Molecular AI's announcement showed, and a $40 million equity stake, BioSpace's data shows. Incyte will also fund recurring research to cover model training and inference compute. In return, Genesis will use Incyte's data, securely, to train the next generation of GEMS, its platform for protein-ligand structure and property prediction. This arrangement is among the first major pharma-AI alliances explicitly designed to support large-scale foundation model training using a partner's proprietary experimental data.
The deal builds on a February 2025 pact with a $30 million upfront fee covering two targets, the companies' filings put, and an option for a third. The expansion adds at least five new targets selected by Incyte, with options to nominate more. Incyte retains exclusive development and commercialization rights to resulting products. Genesis becomes eligible for up to $232 million in preclinical, clinical, regulatory, and sales milestones per program, FierceBiotech found, more than $1 billion across the first five targets, Incyte's release noted, if aggregate peak annual net sales clear specified thresholds. Additional targets could push total potential value several billion dollars higher. Royalties on approved products complete the structure.
"This expanded collaboration reflects the strong results from our initial programs and Incyte's commitment to applying advanced technologies to enhance our discovery engine and accelerate the development of differentiated small molecule medicines," said Pablo J. Cagnoni, M.D., President and Head of Research and Development at Incyte. "By combining our deep expertise in drug discovery and development and our significant experimental data with Genesis' AI capabilities, we aim to more efficiently advance priority programs against high-value targets and ultimately bring important new medicines to patients."
Genesis, founded in 2019 by Evan Feinberg out of Vijay Pande's Stanford lab, raised $200 million in 2023, Genesis's news page reported, and signed a partnership with Gilead Sciences in 2024. Its GEMS platform integrates generative and predictive AI with physics-based simulation, including Pearl, a diffusion model for 3D structure prediction. Chemists use the platform to create candidates, predict properties, interrogate predictions, and decide what to synthesize, with AI agents orchestrating the cycle.
"Our partnership with Incyte reflects the exciting synergy between our AI platform and Incyte's deep expertise and capabilities in rapid experimental data generation and marks an important moment in the evolution of AI in this vertical," said Feinberg. "High-quality proprietary data is among the most valuable inputs for advancing molecular AI, and our expanded collaboration will enable both companies and patients to benefit from an industrial-scale flywheel of AI-enabled design-make-test cycles."
Ropes & Gray advised Genesis on the transaction. Bristol Myers Squibb, meanwhile, unveiled a separate agreement with Anthropic the same week to use Claude across discovery, development, and delivery — a reminder that the industry is testing multiple models for the same shift.
The Old Script Is Broken
For years, pharma-AI partnerships followed a predictable script: a drug company paid an AI vendor for platform access, pointed it at a target or two, and hoped for a candidate. The AI company kept its model weights; the pharma company kept its data. The relationship was transactional — software as a service, not science as a partnership.
The Incyte-Genesis agreement breaks that pattern. Incyte is not licensing Genesis' GEMS platform. It is feeding it. Under the expanded deal, Incyte will securely share proprietary experimental data to train Genesis' next-generation foundation models. Feinberg put it bluntly: the partnership marks that milestone, calling proprietary data the key input for advancing molecular AI and describing the resulting flywheel of AI-enabled design-make-test cycles.
That flywheel is the structural novelty. In the old model, each program started from a cold start — the AI had no memory of the pharma partner's past failures, its chemical series, its target-class quirks. In the new model, every experiment Incyte runs enriches the model that will design the next compound. The data becomes a compounding asset, not a one-time input.
Intuition Labs, which tracks deal structures across the sector, classifies this as a "data-licensing and model-training agreement" — distinct from traditional R&D collaborations focused on co-developing a specific drug. Payment terms reflect the shift: substantial upfront fees, milestone payments tied to development outcomes, and recurring research funding to cover compute for model training and inference. The pharma partner retains rights to any resulting products; the AI partner gets a model that improves with every new dataset.
Public datasets (ChEMBL, PubChem, PDB) are necessary but insufficient. The edge lives in the wet-lab results that never get published — the failed series, the selectivity cliffs, the PK surprises. This structure is spreading. In January 2026 alone, Eli Lilly partnered with Chai Discovery to build an exclusive AI model trained on Lilly's proprietary biologics data; GSK licensed Noetik's cancer "virtual cell" foundation models for $50 million upfront; Pfizer deployed Boltz's technology across small-molecule programs. AstraZeneca committed $200 million to build a multimodal oncology foundation model on Tempus's 7.3 million-patient dataset. Recursion Pharmaceuticals agreed to pay up to $160 million over five years for access to Tempus's de-identified clinico-molecular database plus its analytics platform. Merck struck a multi-year R&D partnership with Mayo Clinic to integrate clinical, genomic, and imaging data via the Mayo Clinic Platform.
Each deal varies in scope: some license patient data, others experimental chemistry data, others clinical outcomes, but the architecture converges: pharma contributes the data moat; the AI partner contributes the model architecture and compute; both share in the resulting IP. Intuition Labs describes it as a "data-as-subscription model" where pharma treats data as a licensed service. The same analysis notes a "supplier-centric model" that "empowers data generators (startups, labs) to maintain control over their datasets while monetizing them."
The shift redefines what pharma buys. It is no longer buying predictions; it is buying a model that learns its chemistry. The competitive question changes from "which platform performs best on benchmarks?" to "which partner has the data that makes the model work for my targets?"
Three Models, One Market
The Incyte-Genesis deal landed in a market that has already sorted itself into three viable models: tools-as-software, pipeline-with-platform, and partnership-only. The $120 million expansion, specifically its structure around proprietary data training a foundation model, validates the pipeline-with-platform approach while pressuring every player to clarify what they own versus what they license.
Insilico Medicine sits furthest ahead on clinical proof. The Hong Kong-listed company raised HKD 2.277 billion in the largest biotech IPO of 2025, reaching a market capitalization of roughly $2.7 billion. As of April 2026 it had nominated 30 preclinical candidates and secured IND clearance for 13 programs across fibrosis, oncology, immunology, and CNS, with three Phase II trials underway. CEO Alex Zhavoronkov said in mid-2025 the company had advanced 22 developmental candidates, 10 of them clinical-stage. The standout remains rentosertib, a TNIK inhibitor for idiopathic pulmonary fibrosis where both target and molecule were AI-generated. Phase 2a data published in Nature Medicine showed a mean forced vital capacity improvement of +98.4 mL versus a -20.3 mL decline on placebo over 12 weeks. Insilico registered a 320-patient Phase III study (NCT07687459) on July 7, 2026, with an estimated start date of August 30. Three weeks later the FDA granted Fast Track Designation to ISM6331, a pan-TEAD inhibitor for advanced mesothelioma — the company's first such designation. A separate NLRP3 inflammasome inhibitor, ISM8969, partnered with Hygtia Therapeutics for up to $66 million in upfront and milestone payments, completed first-in-human dosing in June 2026. Insilico projected first-half 2026 revenue of $102.5–106.5 million, up 273–287 percent year over year.
Recursion Pharmaceuticals took a different route to scale: acquisition. The $688 million all-stock purchase of Exscientia closed November 20, 2024, merging Recursion's industrial-scale phenotypic screening (millions of cell-painting images weekly) with Exscientia's lead-optimization engine and a Bristol Myers Squibb partnership that had generated $1.2 billion in milestone payments before BMS discontinued the PKC-theta asset in October 2025. The combined pipeline includes REC-4881 (Phase 2, MEK1/2 inhibitor for familial adenomatous polyposis, 43–53 percent polyp burden reduction), REC-617 (CDK7 inhibitor, confirmed partial response in platinum-resistant ovarian cancer), REC-1245 (RBM39 degrader, discovery-to-candidate in 18 months), REC-3565 (MALT1 inhibitor, MHRA Phase 1 clearance January 2025), and the legacy Exscientia asset EXS74539/REC-4539 (LSD1 inhibitor, new Phase 1 ENLYGHT trial launched April 13, 2026). But the merger also forced hard choices: in May 2025 Recursion discontinued three clinical programs (REC-2282, REC-994, REC-3964) and paused REC-4539, despite REC-994's Phase 2 SYCAMORE readout showing 50 percent of high-dose patients with reduced lesion volume versus 28 percent on placebo. Shares dropped nearly 17 percent to $4.76 the next trading day. Q1 2026 results showed a $117.5 million net loss against $665.2 million in cash. The company cut roughly 20 percent of staff to extend runway into 2028. Partnerships with Roche/Genentech ($150 million upfront, up to $12 billion across 40 programs) and Bayer (up to $1.5 billion in milestones) remain intact, and Nvidia's $50 million PIPE from July 2023 brought cloud-computing access for foundation-model work.
Isomorphic Labs, the Alphabet spinout built on AlphaFold, has raised the most capital with the least clinical evidence. Same-day collaborations with Eli Lilly ($45 million upfront, up to $1.7 billion milestones) and Novartis ($37.5 million upfront, up to $1.2 billion milestones) in January 2024 totaled "nearly $3 billion" excluding royalties; Novartis added up to three more programs in February 2025. A $600 million Series A led by Thrive Capital closed March 2025, followed by a $2.1 billion Series B in May 2026 — the second-largest biotech round ever, behind only Altos Labs. AlphaFold3, released May 2024, predicts full biomolecular complexes including protein-ligand and protein-nucleic-acid structures, a step change for structure-based design. Yet as of mid-2026 Isomorphic had not disclosed a named clinical candidate or dosed a patient. Demis Hassabis told the World Economic Forum in January 2026 he expected first clinical trials by year-end, revising earlier guidance for end of 2025. President Colin Murdoch said in July 2025 the company was "getting ready to start testing its AI-designed drugs in humans." The internal pipeline is described as concentrated in oncology and immunology.
| Company | Clinical Programs (Phase 2+) | Cash / Runway | Key Partnerships | Latest Catalyst |
|---|---|---|---|---|
| Insilico Medicine | 3 Phase II, 1 Phase III registered | ~$102–106M H1 2026 revenue | Sanofi, Hygtia ($66M) | Rentosertib Phase III start est. Aug 2026; ISM6331 Fast Track |
| Recursion + Exscientia | 5+ Phase 1/2 (REC-4881, REC-617, REC-1245, REC-3565, REC-4539) | $665M cash (Q1 2026) | Roche/Genentech ($12B), Bayer ($1.5B), Nvidia | REC-1245 Phase 1 data H1 2026; 20% staff cut |
| Isomorphic Labs | 0 (pre-clinical) | $2.7B+ raised (Series A+B) | Lilly ($1.7B), Novartis ($1.2B+) | AlphaFold3 integration; first-in-human target end 2026 |
| Generate Biomedicines | 1 Phase III (GB-0895, 2 trials, ~1,600 pts) | $400M IPO Feb 2026 | Amgen | SOLAIRIA-1/2 enrollment Dec 2025 |
| Xaira Therapeutics | 0 (pre-clinical) | >$1B committed (ARCH, Foresite) | Undisclosed | Antibody focus; no candidate disclosed |
| Iambic Therapeutics | 1 Phase 1/1b (IAM1363, 28% ORR HER2+) | Undisclosed | Undisclosed | Goal: 3 drugs in clinic 2026 |
Generate Biomedicines priced its IPO at $16 per share on February 26, 2026, raising $400 million — the year's largest biotech offering to that point. The S-1 disclosed a 2025 net loss of about $223 million against roughly $32 million in collaboration revenue. Its lead candidate GB-0895, a long-acting anti-TSLP antibody dosed once every six months, entered two global Phase 3 trials (SOLAIRIA-1 and SOLAIRIA-2) enrolling approximately 1,600 severe asthma patients in December 2025, with a COPD Phase 1b readout expected later in 2026. An earlier program, GB-0669, moved "from computer to clinic in just 17 months." Xaira launched in April 2024 with more than $1 billion in committed capital from ARCH Venture Partners and Foresite Labs, the largest initial funding commitment in ARCH's history, but has disclosed no clinical candidate, focusing on antibody therapeutics in immunology and inflammation. Iambic Therapeutics' HER2 inhibitor IAM1363 showed partial responses in 28 percent of heavily pretreated patients, including those previously treated with trastuzumab deruxtecan and tucatinib, per October 2025 ESMO data. The company called it "one of the first clear demonstrations of a drug candidate with compelling clinical activity and safety from a TechBio company," with program start to clinical trial initiation in two years. The Phase 1/1b trial remained active and recruiting as of early 2026, and Iambic has stated a goal of three drugs in the clinic during 2026.
Schrödinger occupies the tools-as-software lane, selling physics-based simulation (FEP+, Glide, Maestro) to nearly every large pharma while advancing its own pipeline. Its MALT1 inhibitor SGR-1505 achieved a 22 percent overall response rate in relapsed/refractory B-cell lymphoma with dose escalation complete in June 2025. A second program, the CDC7 inhibitor SGR-2921, was halted in August 2025 after two treatment-related deaths in Phase 1. The argument for physics-plus-ML remains that physics reliably predicts binding affinity on novel chemotypes while ML triages billion-compound libraries down to the thousand physics evaluates rigorously. Smaller platforms (BenevolentAI (restructuring after BEN-2293's Phase 2 failure), Atomwise (pivot to partnerships), Verge Genomics (shifted to partnership model after VRG50635 missed its Phase 1b endpoint)) have largely exited the pipeline-with-platform tier.
The next 18 months will be decided in the clinic: whether rentosertib confirms in a larger trial, whether Isomorphic's first molecules dose patients on schedule, and whether any AI-originated asset reaches a Phase 3 readout. The companies that survive to their next readout, with cash and focus intact, will write the next chapter.
The Data Moat
The Incyte-Genesis deal exposes a shift building for two years: the model is no longer the product. The data is.
Every AI platform in drug discovery now runs on similar architectures: graph neural networks, transformer variants, diffusion models for molecular generation. The compute is rented from the same cloud providers. The talent pools overlap. What separates a platform that delivers clinical candidates from one that generates pretty structures is the training set. And the training sets that matter — millions of experimental binding affinities, cellular phenotypic screens, ADMET profiles, and crucially, the negative data that never gets published — sit inside pharma's firewalls.
Public databases like ChEMBL and PubChem cover only a fraction of the chemical space pharma has actually explored. The other 99% — the failed compounds, the toxicity liabilities, the solubility nightmares — is proprietary. That negative space is where foundation models learn what not to generate.
Incyte's expanded agreement with Genesis makes this explicit. The $120 million upfront (including $40 million equity) buys more than access to the GEMS platform. It funds the ingestion of Incyte's experimental data into Genesis's next-generation foundation model. Incyte becomes a co-author of the model weights. The same logic drove Roche's $150 million upfront partnership with Recursion, Sanofi's $100 million expansion with Exscientia, and Bayer's $80 million rare-disease deal with Recursion. Pharma has shifted from purchasing predictions to paying to embed its institutional memory into the prediction engine.
The competitive map now sorts cleanly by data gravity.
| Player Type | Data Assets | Clinical Programs | Revenue Tier | Moat Durability |
|---|---|---|---|---|
| Tier 1 Integrated Platforms (Insilico, Recursion-Exscientia, Relay, Schrödinger) | Proprietary wet-lab data + partner pharma data | 3–10 | $50–200M | High — self-generating data loops |
| Big Pharma Internal (Pfizer, Roche, Novartis, AstraZeneca, Sanofi) | Decades of experimental, clinical, manufacturing data | 100s | $7–9B annual AI R&D | Highest — but siloed, legacy formats |
| Tier 2 Specialized Tech (Atomwise, BenevolentAI, Insitro, Generate, Iktos) | Niche assay data, focused chemical series | 0–3 | $10–50M | Medium — dependent on pharma partnerships |
| Tier 3 Point Solutions | Public data + limited proprietary | Rare | $1–10M | Low — commoditized |
The 50:1 ratio between announced 'biobucks' and actual upfront payments reveals appropriate industry caution. Pharma pays for data access, not promises.
Tier 1 platforms have built the only scalable alternative to pharma's internal data: their own automated wet labs. Recursion runs millions of experiments weekly in its Salt Lake City facility. Insilico's robotic labs generate comparable throughput. These companies don't just license models — they manufacture the training data. That vertical integration is why they command 8–12x revenue multiples for clinical-stage assets.
Big pharma holds deeper data but struggles to use it. Pfizer spends $800 million annually on AI; Roche has a $500 million Basel AI center; Novartis commits $400 million to generative chemistry. Yet 68% of tech executives cite poor data quality and governance as the primary reason AI initiatives fail. The data exists in fragmented LIMS, ELN, and CDMS systems, annotated inconsistently across decades, often missing the metadata (cell passage number, reagent lot, operator) that makes it trainable. The winners among big pharma will be those who solve data engineering first — not those who hire the most ML researchers.
Smaller AI companies face a tightening vise. Without wet-lab infrastructure, they depend on pharma partnerships for training data. But the Incyte-Genesis structure shows pharma now demands equity and model co-ownership in return. The "platform deal" of 2021–2023, $10–20 million upfront for target-specific access, is dying. In its place: foundation-model partnerships where pharma contributes data, compute, and capital in exchange for model weights and IP rights. Companies that cannot offer either proprietary data generation or a path to model co-ownership will be acquired for talent or wound down. The 2025 shakeout, multiple shutdowns, 20%+ workforce cuts, delistings, was the first wave. The second wave hits any Tier 2 or 3 company without a signed foundation-model deal by end of 2026.
The moat is not the algorithm. It is the curated, annotated, negative-data-rich experimental corpus that the algorithm feeds on. Incyte just proved it will pay nine figures to turn its corpus into someone else's model weights. Every other pharma company with a patent cliff approaching is running the same calculation.
What the FDA Now Requires
The FDA's January 2025 draft guidance, "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision Making for Drug and Biological Products," marked the agency's first comprehensive attempt to apply a coherent regulatory standard to AI-derived evidence. The guidance centers on a risk-based credibility assessment framework built around a single organizing concept: context of use. An AI model's credibility is not intrinsic — it is earned for a specific role, whether predicting ADMET properties, selecting clinical trial patients, or designing molecules. The framework defines three risk tiers. High risk: the AI directly determines patient safety, drug quality, or pivotal trial outcomes. Medium risk: the AI informs decisions but human oversight and validation are present. Low risk: the AI supports exploratory or hypothesis-generating activities. For each tier, the guidance maps five credibility domains (data quality and governance, model architecture and development, model performance and validation, bias mitigation and fairness, lifecycle management and continuous monitoring) and prescribes a seven-step process to establish and assess credibility commensurate with the model's influence and the consequence of a wrong decision.
This structure has immediate consequences for partnerships like Incyte-Genesis. When Incyte's proprietary experimental data trains Genesis's next-generation GEMS platform, that training data becomes part of the regulatory evidence chain. The guidance makes clear that data quality and governance are foundational: variability in the quality, size, and representativeness of training datasets may introduce bias and raise questions about reliability. For a foundation model trained on a pharma partner's private data, the sponsor must document provenance, curation standards, and representativeness, not just for the model developer's internal records, but for FDA review. The guidance also flags data drift: AI-based models may be highly sensitive to variations in model inputs because they are data-driven and can be self-evolving, capable of autonomously adapting without human intervention. Performance metrics must be monitored on an ongoing basis to ensure the model remains fit for use over the drug product life cycle for its context of use. That lifecycle obligation falls on the sponsor, not the AI vendor alone.
The agency's experience is scaling fast. Since 2016, the use of AI in drug development and regulatory submissions has increased exponentially. CDER reports more than 500 drug and biologic submissions with AI components from 2016 through 2023, with submissions growing roughly 40 percent year over year from 2021 to 2023. An internal CDER audit from November 2025 found that 60 percent of NDAs in 2024 contained at least one AI-generated analysis, such as population modeling. About 10 percent of IND applications now include trial protocols with remote monitoring plans. The acceptance rate is high: over 95 percent of AI-supported submissions are not rejected solely due to AI use, but common issues requiring clarification cluster in five areas: insufficient validation data (35 percent of submissions), unclear context of use (28 percent), data quality concerns (18 percent), inadequate bias assessment (12 percent), and other (7 percent). The pattern is clear: sponsors who treat AI as a back-office tool, separate from regulatory strategy, face longer review cycles. Those who engage early, disclosing AI use in pre-IND meetings, classifying each use case by regulatory risk, developing credibility assessment templates, report shorter reviews.
The regulatory calendar is accelerating. In January 2026, the FDA and EMA jointly published the "Guiding Principles of Good Machine Learning Practice for Drug Development," ten high-level principles designed to align international expectations and complement the credibility framework. The CDER AI Council, established in 2024, now coordinates the center's AI policy, internal capabilities, and external communications. CDER's 2026 guidance agenda, released February 2026, lists two draft guidances directly relevant: "Use of Digital Health Technologies in Clinical Investigations of Drugs and Biologics" and "AI/ML Quality Considerations in Pharmaceutical Manufacturing." The latter will likely address validating and verifying AI models, establishing performance metrics, ensuring data integrity, documenting controls, and may cover supply-chain aspects of third-party AI tools or cloud platforms. A final version of the AI credibility guidance is expected in Q2 2026, incorporating comment-period revisions and alignment with the joint FDA/EMA principles. In parallel, the FDA launched an AI-Enabled Optimization of Early-Phase Clinical Trials pilot program in April 2026, offering selected sponsors direct technical engagement on adaptive randomization algorithms, dose-finding models, and AI-augmented patient selection. Lessons from the pilot will inform a subsequent guidance on AI-driven trial design. A real-time clinical trial data monitoring pilot, also announced in April 2026, aims to use cloud infrastructure and AI to monitor trial data in near-real-time; FDA officials have said the program could ultimately reduce clinical trial timelines by 20 to 40 percent.
The first AI-discovered drug approval is anticipated in 2026-2027, with roughly 60 percent probability. Insilico Medicine published Phase IIa results for rentosertib (ISM001-055) in Nature Medicine in June 2025 — the first such publication for a fully AI-discovered drug. Isomorphic Labs cleared its first AI-designed molecule, ISM8969, for human trials in January 2026. Recursion has logged more than $500 million in upfront payments from partnerships. The clinical signal is arriving. For Incyte and Genesis, the regulatory horizon means their data-sharing architecture must be auditable from day one. The credibility of a foundation model trained on Incyte's proprietary assays will be judged not only on predictive performance but on the documented chain from raw experimental data through curation, training, validation, bias assessment, and lifecycle monitoring. The FDA has signaled that it will not accept a black box, even a well-performing one, when the model's outputs influence regulatory decisions. The partnership's value, ultimately, will be measured in the regulatory currency of credibility evidence, and the sponsors who build that evidence into their collaboration agreements now will file INDs with fewer surprises later.
The Talent Gap No One Can Hire Around
The numbers are stark: 1.6 million open AI-related positions globally against 518,000 qualified candidates — a supply deficit exceeding three to one. Job postings requiring AI skills have jumped 247 percent since 2023 while the candidate pool grew just 63 percent. For the first time in the survey's history, AI fluency has overtaken engineering as the hardest competency for employers to find worldwide.
This shortage hits biotech differently than it hits fintech or enterprise software. The Incyte-Genesis partnership, where a pharmaceutical company contributes proprietary experimental data to train a foundation model, creates a new class of hiring demand. The model only works if someone understands how the data was generated, why the assay conditions matter, and what the regulatory reviewer will ask about provenance. That person is neither a traditional data scientist nor a traditional biologist. They are the "bilingual" expert who can translate a scientific problem into something an AI team can actually build, then evaluate whether the output holds up under GxP scrutiny.
Biotech has always competed for specialized talent. A gene therapy company cannot simply hire a strong scientist; it needs someone with deep expertise in the specific science behind its work. Now that requirement has doubled. A software engineer can build an impressive machine-learning system and still miss something scientifically important about the underlying data. A bench scientist can understand the biology extremely well but may not know how to evaluate data leakage, model performance, infrastructure constraints, or whether a system can be reproduced reliably. Biological data presents its own problems: noisy, sparse, multimodal, expensive to generate, with ground truth that is not always obvious. Reproducibility and data provenance matter just as much as model accuracy.
The competition for these hybrids is fierce. Tech giants bid for the same limited pool, and the largest waves of AI hiring now come from sectors applying AI to their own problems, meaning industry knowledge has become as valuable as technical skill. U.S. postings for AI, machine learning, and data science roles jumped 163 percent from 2024 to 2025. LinkedIn ranked AI engineer the fastest-growing job title in the country heading into 2026. Workers with advanced AI skills command a 56 percent wage premium over peers in identical roles without those skills, a premium that nearly doubled in a single year.
Companies are responding, but the response is uneven. Eighty-two percent of enterprises provide AI training, yet 59 percent still report a significant skills gap. Only 29 percent of workers globally say their workplace invests enough in AI training. Of employees who received training, 65 percent learned primarily through self-study; formal employer programs accounted for just 24 percent. Only 18 percent say the training prepared them to work independently. Meanwhile, 67 percent of employees say they need two hours or fewer per week to meaningfully improve their AI skills, yet even that time is rarely protected.
The retention math is brutal. AI heavy users, the most skilled employees, are 7 to 10 percentage points more likely than light users to plan on quitting in the next three to six months, because they know their skills are in high demand. Thirty-five percent of employees say they will look for a new job if their company does not provide adequate training, up from 24 percent a year prior. Companies that invest in AI training spend an average of $2,100 per employee; those that do not spend $380 and see retention rates 12 percentage points lower at 24 months.
Some organizations are shifting strategy. Two-thirds of companies say they plan to hire for specific AI skills rather than generic ones. The throughline across sectors is the same: the most valuable AI professionals are no longer just the ones who can build a model in isolation. They are the ones who can apply AI to a specific domain, deploy it responsibly, and explain why it matters to the business. Smart biotech firms are looking inward first: a computational biologist already on staff who develops deeper machine-learning skills may ultimately be more useful than a newly hired generic AI specialist who needs years to develop domain knowledge. The same holds for scientists who learn to design AI-enabled workflows or engineers who develop a serious understanding of the biology behind the systems they are building.
The junior pipeline is compressing. Fifty-one percent of organizations report that generative AI is reducing their need for entry-level roles, cutting off the traditional skill-building path for early-career workers. AI-exposed junior roles are seven times more likely to require traditionally senior skills such as leadership and strategic thinking compared to roles with low AI exposure. Sixty-five percent of organizations have abandoned AI projects due to skills gaps — direct write-downs on technology investment.
By the end of 2026, more than 90 percent of global organizations will experience AI skills shortages, according to IDC projections. The gap represents an estimated $5.5 trillion in unrealized global productivity. Incyte's $120 million just bought a seat at the table where the next generation of foundation models gets written — with pharma's proprietary failures as the ink. The companies that understand data is the product, not the platform, will own the next decade of drug discovery.
Working in frontier tech? Zero G Talent tracks the openings: see every open ASML role, browse frontier tech jobs, openings at Stripe, and the people building the field.