In partnership with

AI Spotlight — The Real Bottleneck in AI Drug Discovery Isn't the Model
AI SPOTLIGHT

The Real Bottleneck in AI Drug Discovery Isn't the Model

GSK just paid up to $110 million to generate more cells, not more compute. Here's why that's the smarter bet.

📖 5 minute read
Laboratory researcher examining cell samples under a microscope

Welcome Back,

Most AI headlines chase bigger models. This one is about something less glamorous and arguably more important, better biological data.

GSK has entered a research collaboration with British biotech Relation Therapeutics worth up to $110 million, expanding work the two companies were already doing together on AI-assisted drug discovery, according to AI News. Under the deal, Relation will generate large-scale datasets measuring how human cells actually respond to genetic changes and drug interventions, then use that data to train AI models built to identify potential drug targets.

Today we look at exactly what Relation's data-generation process involves, why public biological datasets have real limits researchers are actively working around, a genuinely counterintuitive research finding about model scale, and what this deal fits into as a broader industry pattern.

📌 In Today's AI Spotlight

  • What's actually in the GSK-Relation Therapeutics deal.
  • How Relation's "Lab-in-the-Loop" approach generates biological data.
  • Why public single-cell repositories have quiet, technical limitations.
  • The surprising research finding that more data doesn't always mean better models.
  • Our AI Spotlight take on what this signals for the AI drug discovery industry.

🧬 What's Actually In the Deal

The agreement places biological data generation directly alongside AI model development, rather than treating them as separate workstreams. Relation's research approach links computational analysis with experiments that generate genuinely new information about human cells, not just fresh ways to analyze data that already exists.

This builds on earlier agreements between the two companies focused specifically on fibrotic diseases and osteoarthritis. Those projects involved observational studies designed to create two functional disease datasets, analyzed using Relation's Lab-in-the-Loop platform, combining human genetics, single-cell multi-omics from human tissue, functional assays, and machine learning to identify and validate potential disease targets.

Relation will produce human cellular datasets as part of the collaboration and use them to train AI models for identifying potential drug targets, including models within Relation's MORGAN platform.

That combination, generating the biology and training the model in the same pipeline, is the part worth paying attention to. It's a meaningfully different bet than simply licensing an existing dataset or buying compute to run bigger training runs on data that already exists.

Scientist working with pipettes and lab equipment in a research facility

The deal pairs fresh laboratory experimentation with the AI models trained on its results.

🔬 How Relation Actually Generates Its Data

Relation describes its Lab-in-the-Loop approach as a genuine combination of laboratory experimentation and computational analysis, not one feeding into the other in a single direction. Its work spans tissue profiling, single-cell and spatial transcriptomics, sequencing, and target validation, with machine learning used throughout for target identification, prioritization, validation, and experimental design.

💡 AI Spotlight Take

The phrase "in the loop" is doing real work here. Most AI pipelines treat data collection and model training as sequential, collect first, train second. Relation's approach feeds model outputs back into deciding what experiment to run next, which means the AI is actively shaping what biological data gets generated, not just consuming whatever's already sitting in a database.

The company also conducts perturbation experiments, deliberately introducing genetic changes and measuring how they affect cellular characteristics associated with disease. Those results then get analyzed alongside genetic and patient-derived biological data, building a dataset that's purpose-built for the specific disease question at hand, rather than repurposed from a general-use public archive.

Thinking about hiring globally? Start with an EOR.

The best person for your next role might not live near your office—or even in the same country.

More companies are realizing they don't need to open entities everywhere just to access global talent. Instead, they're using EOR to hire internationally faster, stay compliant, and avoid building local infrastructure before they're ready.

Oyster's EOR helps companies hire, pay, and support employees in 180+ countries while Oyster handles payroll, compliance, taxes, and local employment requirements.

AI Spotlight — The Real Bottleneck in AI Drug Discovery Isn't the Model Part 2

📚 Why Public Datasets Aren't Enough on Their Own

Public repositories remain genuinely important for training biological foundation models. A 2025 review in Experimental & Molecular Medicine noted that repositories including CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus give researchers access to enormous volumes of single-cell data, CZ CELLxGENE alone provides more than 100 million standardized cells.

The Deal By the Numbers

$110M

maximum value of the GSK-Relation collaboration

 

100M+

standardized cells available through CZ CELLxGENE alone

 

22.2M

cells used in a landmark study on model scaling limits

The catch is that sampling methods, sequencing protocols, and processing pipelines differ between studies, and single-cell data can carry technical noise and other artefacts, requiring careful dataset selection, filtering, and quality control before it's usable for training. Dataset overlap adds another wrinkle, the same or similar cells can show up across multiple public resources, giving them outsized influence during training and creating data-leakage risk when training and test sets accidentally overlap.

Researcher analyzing genomic data visualizations on a computer screen

Assembling a clean, non-redundant dataset turns out to matter as much as the model architecture built on top of it.

📉 Bigger Datasets Don't Automatically Mean Better Models

Here's the finding that should reshape how anyone thinks about this space. Research published in Nature Methods this past June examined how the size and diversity of pretraining data affected single-cell foundation models, using a corpus of 22.2 million cells, training 400 models and evaluating them across 6,400 experiments.

Current single-cell foundation models tended to reach performance plateaus after training on only a fraction of the available corpus. Unlike large language models, these systems did not display clear data-scaling laws where continually increasing training data consistently produced better results.

That's a genuinely important departure from how most people think AI progress works, more data equals better models. The researchers found that model capacity, dataset size, and computational resources need to be balanced against each other, not simply increased together, though the study stopped short of claiming smaller or proprietary datasets are inherently better.

A separate 2025 study in Genome Biology reinforces the point. It assessed two single-cell foundation models, Geneformer and scGPT, and found they didn't consistently outperform simpler approaches across several zero-shot evaluation tasks, while flagging challenges around batch effects and cautioning against assuming larger pretrained models automatically produce better biological representations.

🏭 A Broader Pharma Industry Pattern

GSK and Relation aren't operating in isolation here. Relation has already applied its data-generation approach to Osteomics, a proprietary functional single-cell bone atlas combining single-cell and spatial omics with imaging, genomics, proteomics, and clinical phenotype data from patient-derived samples, with hospitals and research partners in the UK and Australia involved in the observational study.

Other Specialized Biological Dataset Deals

⚠️  GSK's separate $37.5M agreement with Ochre Bio for human liver single-cell and perfused-organ data
⚠️  AstraZeneca and Pathos AI's $200M 2025 agreement with Tempus, covering de-identified data from more than 150,000 oncology patients
⚠️  A 2025 Nature Biotechnology analysis identified specialized dataset providers as one of several emerging trends in AI-biopharma deals

That analysis also flagged larger upfront payments, new therapeutic modalities, and greater participation from bigger biotech companies as related trends, all pointing in the same direction, high-quality, disease-specific datasets are becoming a genuinely important input for causal and generative machine-learning models, not a secondary consideration behind the AI itself.

Pharmaceutical research team reviewing molecular structures on a screen

Access to sufficient high-quality data remains one of the central bottlenecks in AI drug discovery.

🧠 AI Spotlight Analysis

There's a genuine tension worth naming here. The industry keeps signing deals premised on the idea that more specialized biological data leads to better drug-target models, while the actual research on single-cell foundation models keeps finding that scale alone doesn't reliably deliver that improvement. Both things can be true at once, quality and relevance may matter more than raw volume, which is precisely the bet a purpose-built, disease-specific dataset like Relation's is making.

A Nature research highlight on federated learning in pharmaceutical research backs up why deals like this keep happening, limited access to suitable training data remains a major bottleneck for AI applications in this field, compounded by the restrictions companies often face on sharing proprietary information at all.

💬 Quote of the Week

Assembling a high-quality, non-redundant dataset is as important as model architecture when building robust single-cell foundation models.

— 2025 review, Experimental & Molecular Medicine

That line captures the real shift underway in AI drug discovery. The story isn't just about which lab has the sharpest model, it's increasingly about which lab has generated data clean and specific enough for that model to actually learn something true.

💡 Final Thoughts

The GSK-Relation deal is a useful reminder that the most consequential AI stories aren't always about a new model release. Sometimes the more important move is paying up to $110 million to generate cleaner, more specific biological data, precisely because the research now shows that scaling data volume alone doesn't reliably improve these models the way it does with language.

As more pharma companies chase specialized, disease-specific datasets rather than simply bigger public archives, the real competitive edge in AI drug discovery may increasingly belong to whoever can generate the right biological data, not just whoever has the biggest model.

Does it surprise you that more biological data doesn't always mean a better AI model? Hit reply, we read every response.

🔗 Sources and Further Reading

AI News: Why biological data matters more in AI drug discovery

❤️ Enjoying AI Spotlight?

If today's edition helped you rethink what actually drives progress in AI drug discovery, consider sharing it with a colleague, founder, or friend interested in technology.

Share AI Spotlight →

Thanks for reading AI Spotlight.

Our mission is simple: deliver clear, trustworthy, and actionable AI insights that help professionals stay ahead without the hype.

Keep Reading