📚 Why Public Datasets Aren't Enough on Their Own
Public repositories remain genuinely important for training biological foundation models. A 2025 review in Experimental & Molecular Medicine noted that repositories including CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus give researchers access to enormous volumes of single-cell data, CZ CELLxGENE alone provides more than 100 million standardized cells.
|
The Deal By the Numbers
|
$110M
maximum value of the GSK-Relation collaboration
|
|
100M+
standardized cells available through CZ CELLxGENE alone
|
|
22.2M
cells used in a landmark study on model scaling limits
|
|
The catch is that sampling methods, sequencing protocols, and processing pipelines differ between studies, and single-cell data can carry technical noise and other artefacts, requiring careful dataset selection, filtering, and quality control before it's usable for training. Dataset overlap adds another wrinkle, the same or similar cells can show up across multiple public resources, giving them outsized influence during training and creating data-leakage risk when training and test sets accidentally overlap.
|
Assembling a clean, non-redundant dataset turns out to matter as much as the model architecture built on top of it.
|
📉 Bigger Datasets Don't Automatically Mean Better Models
Here's the finding that should reshape how anyone thinks about this space. Research published in Nature Methods this past June examined how the size and diversity of pretraining data affected single-cell foundation models, using a corpus of 22.2 million cells, training 400 models and evaluating them across 6,400 experiments.
Current single-cell foundation models tended to reach performance plateaus after training on only a fraction of the available corpus. Unlike large language models, these systems did not display clear data-scaling laws where continually increasing training data consistently produced better results.
That's a genuinely important departure from how most people think AI progress works, more data equals better models. The researchers found that model capacity, dataset size, and computational resources need to be balanced against each other, not simply increased together, though the study stopped short of claiming smaller or proprietary datasets are inherently better.
A separate 2025 study in Genome Biology reinforces the point. It assessed two single-cell foundation models, Geneformer and scGPT, and found they didn't consistently outperform simpler approaches across several zero-shot evaluation tasks, while flagging challenges around batch effects and cautioning against assuming larger pretrained models automatically produce better biological representations.
|
🏭 A Broader Pharma Industry Pattern
GSK and Relation aren't operating in isolation here. Relation has already applied its data-generation approach to Osteomics, a proprietary functional single-cell bone atlas combining single-cell and spatial omics with imaging, genomics, proteomics, and clinical phenotype data from patient-derived samples, with hospitals and research partners in the UK and Australia involved in the observational study.
|
Other Specialized Biological Dataset Deals
| ⚠️ GSK's separate $37.5M agreement with Ochre Bio for human liver single-cell and perfused-organ data |
| ⚠️ AstraZeneca and Pathos AI's $200M 2025 agreement with Tempus, covering de-identified data from more than 150,000 oncology patients |
| ⚠️ A 2025 Nature Biotechnology analysis identified specialized dataset providers as one of several emerging trends in AI-biopharma deals |
|
That analysis also flagged larger upfront payments, new therapeutic modalities, and greater participation from bigger biotech companies as related trends, all pointing in the same direction, high-quality, disease-specific datasets are becoming a genuinely important input for causal and generative machine-learning models, not a secondary consideration behind the AI itself.
|
Access to sufficient high-quality data remains one of the central bottlenecks in AI drug discovery.
|
🧠 AI Spotlight Analysis
There's a genuine tension worth naming here. The industry keeps signing deals premised on the idea that more specialized biological data leads to better drug-target models, while the actual research on single-cell foundation models keeps finding that scale alone doesn't reliably deliver that improvement. Both things can be true at once, quality and relevance may matter more than raw volume, which is precisely the bet a purpose-built, disease-specific dataset like Relation's is making.
A Nature research highlight on federated learning in pharmaceutical research backs up why deals like this keep happening, limited access to suitable training data remains a major bottleneck for AI applications in this field, compounded by the restrictions companies often face on sharing proprietary information at all.
💬 Quote of the Week
Assembling a high-quality, non-redundant dataset is as important as model architecture when building robust single-cell foundation models.
— 2025 review, Experimental & Molecular Medicine
That line captures the real shift underway in AI drug discovery. The story isn't just about which lab has the sharpest model, it's increasingly about which lab has generated data clean and specific enough for that model to actually learn something true.
|
💡 Final Thoughts
The GSK-Relation deal is a useful reminder that the most consequential AI stories aren't always about a new model release. Sometimes the more important move is paying up to $110 million to generate cleaner, more specific biological data, precisely because the research now shows that scaling data volume alone doesn't reliably improve these models the way it does with language.
As more pharma companies chase specialized, disease-specific datasets rather than simply bigger public archives, the real competitive edge in AI drug discovery may increasingly belong to whoever can generate the right biological data, not just whoever has the biggest model.
Does it surprise you that more biological data doesn't always mean a better AI model? Hit reply, we read every response.
|
🔗 Sources and Further Reading
|
❤️ Enjoying AI Spotlight?
If today's edition helped you rethink what actually drives progress in AI drug discovery, consider sharing it with a colleague, founder, or friend interested in technology.
Share AI Spotlight →
|
|
|
Thanks for reading AI Spotlight.
Our mission is simple: deliver clear, trustworthy, and actionable AI insights that help professionals stay ahead without the hype.
|
|