Benchmarking biomedical foundation models

Wait 5 sec.

Bommasani, R. et al. On the opportunities and risks of foundation models. Preprint at https://doi.org/10.48550/arXiv.2108.07258 (2022).Choi, Y. et al. Defining foundation models for computational science: a call for clarity and rigor. Preprint at https://doi.org/10.48550/arXiv.2505.22904 (2025).Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Passaro, S. et al. Boltz-2: towards accurate and efficient binding affinity prediction. Preprint at bioRxiv https://doi.org/10.1101/2025.06.14.659707 (2025).Baek, M. et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871–876 (2021).Article  CAS  PubMed  PubMed Central  Google Scholar Bunne, C. et al. How to build the virtual cell with artificial intelligence: priorities and opportunities. Cell 187, 7045–7063 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Cui, H. et al. Towards multimodal foundation models in molecular cell biology. Nature 640, 623–633 (2025).Article  CAS  PubMed  Google Scholar Ye, C. et al. Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models. Preprint at bioRxiv https://doi.org/10.1101/2025.09.21.677619 (2025).Avsec, Ž. et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature https://doi.org/10.1038/s41586-025-10014-0 (2026).Brixi, G. et al. Genome modelling and design across all domains of life with Evo 2. Nature https://doi.org/10.1038/s41586-026-10176-5 (2026).Chen, R. J. et al. Towards a general-purpose foundation model for computational pathology. Nat. Med. 30, 850–862 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Shmatko, A. et al. Learning the natural history of human disease with generative transformers. Nature https://doi.org/10.1038/s41586-025-09529-3 (2025).Szabó, C. Unreliable: Bias, Fraud, and the Reproducibility Crisis in Biomedical Research (Columbia University Press, 2025).Ritchie, S. Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth (Metropolitan Books, 2021).National Academies of Sciences, Engineering, and Medicine. Reproducibility and Replicability in Science (National Academies Press, 2020).Kapoor, S. & Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4, 100804 (2023).Article  PubMed  PubMed Central  Google Scholar Wong, A. et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern. Med. 181, 1065–1070 (2021).Article  PubMed  PubMed Central  Google Scholar Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K. & Moult, J. Critical assessment of methods of protein structure prediction (CASP)—round XIV. Proteins 89, 1607–1617 (2021). Demonstrated how rigorous community assessment can establish genuine methodological breakthroughs, with CASP14 providing the independent evidence that AlphaFold2 achieved near-experimental accuracy for many single-protein structure prediction targets.Article  CAS  PubMed  PubMed Central  Google Scholar Stolovitzky, G., Monroe, D. & Califano, A. Dialogue on reverse-engineering assessment and methods: the DREAM of high-throughput pathway inference. Ann. N. Y. Acad. Sci. 1115, 1–22 (2007).Article  PubMed  Google Scholar Saez-Rodriguez, J. et al. Crowdsourcing biomedical research: leveraging communities as innovation engines. Nat. Rev. Genet. 17, 470–486 (2016).Article  CAS  PubMed  PubMed Central  Google Scholar Deng, J. et al. ImageNet: a large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition 248–255 https://doi.org/10.1109/CVPR.2009.5206848 (2009). Credited as a catalyst for the deep learning surge, ImageNet was a large-scale yearly competition for image classification, where one of the first breakthrough deep learning architectures, AlexNet, emerged.Szałata, A. et al. A benchmark for prediction of transcriptomic responses to chemical perturbations across cell types. Adv. Neural Inf. Process. Syst. 37, 20566–20616 (2024).Article  Google Scholar Wognum, C. et al. A call for an industry-led initiative to critically assess machine learning for real-world drug discovery. Nat. Mach. Intell. 6, 1120–1121 (2024).Article  Google Scholar Fraser, J. S. et al. OpenADMET: embracing the avoid-ome to transform drug discovery. Preprint at Preprints.org https://doi.org/10.20944/preprints202512.1202.v1 (2025).Reinke, A. et al. Medical imaging AI competitions lack fairness. Preprint at https://doi.org/10.48550/arXiv.2512.17581 (2025).Norel, R., Rice, J. J. & Stolovitzky, G. The self-assessment trap: can we all be better than average?. Mol. Syst. Biol. 7, 537 (2011).Article  PubMed  PubMed Central  Google Scholar Ahsen, M. E., Vogel, R. & Stolovitzky, G. A. R/PY-SUMMA: an R/Python package for unsupervised ensemble learning for binary classification problems in bioinformatics. J. Comput. Biol. 27, 1337–1340 (2020).Article  CAS  PubMed  Google Scholar Kim, S.-C., Arun, A. S., Ahsen, M. E., Vogel, R. & Stolovitzky, G. The Fermi–Dirac distribution provides a calibrated probabilistic output for binary classifiers. Proc. Natl Acad. Sci. USA 118, e2100761118 (2021).Article  CAS  PubMed  PubMed Central  Google Scholar Wornow, M., Thapa, R., Steinberg, E., Fries, J. A. & Shah, N. H. EHRSHOT: an EHR benchmark for few-shot evaluation of foundation models. Adv. Neural Inf. Process. Syst. 36, 67125–67137 (2023).Article  Google Scholar Pang, C. et al. FoMoH: a clinically meaningful foundation model evaluation for structured electronic health records. Preprint at https://doi.org/10.48550/ARXIV.2505.16941 (2025).Liao, Y. et al. EHR-R1: a reasoning-enhanced foundational language model for electronic health record analysis. Preprint at https://doi.org/10.48550/ARXIV.2510.25628 (2025).Johnson, A. E. W. et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data 10, 1 (2023).Article  CAS  PubMed  PubMed Central  Google Scholar Wornow, M. et al. The shaky foundations of large language models and foundation models for electronic health records. NPJ Digit. Med. 6, 135 (2023).Article  PubMed  PubMed Central  Google Scholar Jong, E. D. de, Marcus, E. & Teuwen, J. Current pathology foundation models are unrobust to medical center differences. Preprint at https://doi.org/10.48550/arXiv.2501.18055 (2025).Vásquez-Venegas, C. et al. Detecting and mitigating the clever hans effect in medical imaging: a scoping review. J. Imaging Inform. Med. 38, 2563–2579 (2025).Article  PubMed  Google Scholar Patel, A. et al. DART-Eval: a comprehensive DNA language model evaluation benchmark on regulatory DNA. Preprint at https://doi.org/10.48550/arXiv.2412.05430 (2026).Vieira, L. C., Lin, S. & Wilke, C. O. Intrinsic dataset features drive mutational effect prediction by protein language models. Preprint at bioRxiv https://doi.org/10.64898/2026.03.08.710389 (2026).Tang, Z., Somia, N., Yu, Y. & Koo, P. K. Evaluating the representational power of pre-trained DNA language models for regulatory genomics. Genome Biol. 26, 203 (2025). Showed that representations from current genomic foundation models provide little to no advantage over standard one-hot sequence encodings across several regulatory genomics tasks, highlighting that generic genome-scale pretraining does not necessarily yield transferable cis-regulatory understanding.Article  PubMed  PubMed Central  Google Scholar Antoniuk, E. R. et al. BOOM: benchmarking out-of-distribution molecular property predictions of machine learning models. In Advances in Neural Information Processing Systems Vol. 38 (IEEE, 2025).Zhang, F. et al. A survey on foundation language models for single-cell biology. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds. W. Che et al.) 528–549 https://doi.org/10.18653/v1/2025.acl-long.26 (Association for Computational Linguistics, 2025).Szałata, A. et al. Transformers in single-cell omics: a review and new perspectives. Nat. Methods 21, 1430–1443 (2024).Article  PubMed  Google Scholar Baek, S., Song, K. & Lee, I. Single-cell foundation models: bringing artificial intelligence into cell biology. Exp. Mol. Med. https://doi.org/10.1038/s12276-025-01547-5 (2025).Xie, F. et al. Overcoming barriers to the wide adoption of single-cell large language models in biomedical research. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02846-y (2025).Csendes, G., Sanz, G., Szalay, K. Z. & Szalai, B. Benchmarking foundation cell models for post-perturbation RNA-seq prediction. BMC Genomics 26, 393 (2025).Article  PubMed  PubMed Central  Google Scholar Wang, Q. et al. scDrugMap: benchmarking large foundation models for drug response prediction. Nat. Commun. 17, 730 (2025).Article  PubMed  PubMed Central  Google Scholar Viñas Torné, R. et al. Systema: a framework for evaluating genetic perturbation response prediction beyond systematic variation. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02777-8 (2025). Questioned established metrics for genetic perturbation-response prediction by showing that they can overestimate performance when models capture systematic differences between perturbed and control cells rather than perturbation-specific effects.Li, L. et al. A systematic comparison of single-cell perturbation response prediction models. Preprint at bioRxiv https://doi.org/10.1101/2024.12.23.630036 (2025).Ovcharenko, O. et al. scSSL-Bench: benchmarking self-supervised learning for single-cell data. In Proc. 42nd Int. Conf. Mach. Learn. (PMLR 267) 47416–47442 (2025).Wu, J. et al. Biology-driven insights into the power of single-cell foundation models. Genome Biol. 26, 334 (2025).Article  PubMed  PubMed Central  Google Scholar Kedzierska, K. Z., Crawford, L., Amini, A. P. & Lu, A. X. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biol. 26, 101 (2025).Article  PubMed  PubMed Central  Google Scholar Ahlmann-Eltze, C., Huber, W. & Anders, S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nat. Methods 22, 1657–1661 (2025). Showed that several deep-learning-based perturbation prediction models, including single-cell foundation models, did not outperform simple linear baselines in predicting the effects of unseen genetic perturbations, highlighting that many perturbation-response benchmarks lacked sufficiently strong baselines.Article  CAS  PubMed  PubMed Central  Google Scholar Bendidi, I. et al. Benchmarking transcriptomics foundation models for perturbation analysis: one PCA still rules them all. Preprint at https://doi.org/10.48550/arXiv.2410.13956 (2024).Hasanaj, E. et al. Multimodal benchmarking of foundation model representations for cellular perturbation response prediction. Preprint at bioRxiv https://doi.org/10.1101/2025.06.26.661186 (2025).Wenteler, A. et al. PertEval-scFM: benchmarking single-cell foundation models for perturbation effect prediction. Preprint at bioRxiv https://doi.org/10.1101/2024.10.02.616248 (2024).Wong, D. R., Hill, A. S. & Moccia, R. Simple controls exceed best deep learning algorithms and reveal foundation model effectiveness for predicting genetic perturbations. Bioinformatics 41, btaf317 (2025).Article  CAS  PubMed  PubMed Central  Google Scholar Boiarsky, R. et al. Deeper evaluation of a single-cell foundation model. Nat. Mach. Intell. 6, 1443–1446 (2024).Article  Google Scholar Liu, X. et al. Evaluating foundation models for in-silico perturbation. Preprint at bioRxiv https://doi.org/10.1101/2025.05.11.653338 (2025).Boylan, J. et al. Single cell foundation models evaluation (scFME) for in-silico perturbation. Preprint at bioRxiv https://doi.org/10.1101/2025.09.22.677811 (2025).Atti, S. & Subramaniam, S. Fundamental limitations of foundation models in single-cell transcriptomics. Preprint at bioRxiv https://doi.org/10.1101/2025.06.26.661767 (2025).Liu, T., Li, K., Wang, Y., Li, H. & Zhao, H. Evaluating the utilities of foundation models in single-cell data analysis. Preprint at bioRxiv https://doi.org/10.1101/2023.09.08.555192 (2024).Dandala, B. et al. BMFM-RNA: an open framework for building and evaluating transcriptomic foundation models. Preprint at https://doi.org/10.48550/arXiv.2506.14861 (2025).Qiu, P. et al. BioLLM: a standardized framework for integrating and benchmarking single-cell foundation models. Patterns 6, 101326 (2025).Article  PubMed  PubMed Central  Google Scholar Theodoris, C. V. Perspectives on benchmarking foundation models for network biology. Quant. Biol. 12, 335–338 (2024).Article  PubMed  PubMed Central  Google Scholar Denisenko, E. et al. Systematic assessment of tissue dissociation and storage biases in single-cell and single-nucleus RNA-seq workflows. Genome Biol. 21, 130 (2020).Article  CAS  PubMed  PubMed Central  Google Scholar Rives, A. et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl Acad. Sci. USA 118, e2016239118 (2021).Article  CAS  PubMed  PubMed Central  Google Scholar He, F. et al. Harnessing the power of single-cell large language models with parameter-efficient fine-tuning using scPEFT. Nat. Mach. Intell. 8, 118–133 https://doi.org/10.1038/s42256-025-01170-z (2026).Li, C. et al. Benchmarking AI models for in silico gene perturbation of cells. Preprint at bioRxiv https://doi.org/10.1101/2024.12.20.629581 (2025).Mejia, G. M. et al. Needles in the haystack: addressing signal dilution improves scRNA-seq perturbation response modeling and evaluation. In Proc. 43rd Int. Conf. Mach. Learn. (2026).Kernfeld, E., Yang, Y., Weinstock, J. S., Battle, A. & Cahan, P. A comparison of computational methods for expression forecasting. Genome Biol. 26, 388 https://doi.org/10.1186/s13059-025-03840-y (2025).Miller, H. E. et al. Deep learning-based genetic perturbation models do outperform uninformative baselines on well-calibrated metrics. Preprint at https://doi.org/10.1101/2025.10.20.683304 (2025).Rautenstrauch, P. & Ohler, U. Shortcomings of silhouette in single-cell integration benchmarking. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02743-4 (2025).Wang, H., Leskovec, J. & Regev, A. Limitations of cell embedding metrics assessed using drifting islands. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02702-z (2025). Showed that commonly used single-cell integration benchmarking metrics for evaluating scRNA-seq data integration can fail to detect embeddings that distort biologically meaningful relationships between cell states, highlighting an important shortcoming of current single-cell embedding evaluation frameworks.Luecken, M. D. et al. Benchmarking atlas-level data integration in single-cell genomics. Nat. Methods 19, 41–50 (2022).Article  CAS  PubMed  Google Scholar Ji, B. et al. CAPTAIN: a multimodal foundation model pretrained on co-assayed single-cell RNA and protein. Nat. Commun. 17, 6161 https://doi.org/10.1038/s41467-026-72882-y (2026).Tejada-Lapuerta, A. et al. Nicheformer: a foundation model for single-cell and spatial omics. Nat. Methods https://doi.org/10.1038/s41592-025-02814-z (2025).Rizvi, S. A. et al. Scaling large language models for next-generation single-cell analysis. Preprint at bioRxiv https://doi.org/10.1101/2025.04.14.648850 (2025).Zhang, K. et al. BiomedGPT: a generalist vision-language foundation model for diverse biomedical tasks. Nat. Med. 30, 3129–3141 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Hou, W. & Ji, Z. Assessing GPT-4 for cell type annotation in single-cell RNA-seq analysis. Nat. Methods 21, 1462–1465 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Chen, Y. & Zou, J. GenePert: Leveraging GenePT embeddings for gene perturbation prediction. Preprint at bioRxiv https://doi.org/10.1101/2024.10.27.620513 (2024).Märtens, K., Martell, M. B., Prada-Medina, C. A. & Donovan-Maiye, R. LangPert: LLM-driven contextual synthesis for unseen perturbation prediction. In ICLR 2025 Workshop on Machine Learning for Genomics Explorations (2025).Dua, R. et al. Clinically grounded agent-based report evaluation: an interpretable metric for radiology report generation. Preprint at https://doi.org/10.48550/arXiv.2508.02808 (2025).Mitchener, L. et al. BixBench: a comprehensive benchmark for LLM-based agents in computational biology. Preprint at https://doi.org/10.48550/arXiv.2503.00096 (2025).Miller, H. E., Greenig, M., Tenmann, B. & Wang, B. BioML-bench: evaluation of AI agents for end-to-end biomedical ML. Preprint at bioRxiv https://doi.org/10.1101/2025.09.01.673319 (2025).Scheurer, J., Balesni, M. & Hobbhahn, M. Large language models can strategically deceive their users when put under pressure. Preprint at https://doi.org/10.48550/arXiv.2311.07590 (2024).Heil, B. et al. Reproducibility standards for machine learning in the life sciences. Nat. Methods 18, 1132–1135 (2021).Article  CAS  PubMed  PubMed Central  Google Scholar Roohani, Y. H. et al. Virtual cell challenge: toward a turing test for the virtual cell. Cell 188, 3370–3374 (2025).Article  CAS  PubMed  Google Scholar Walker, C. R. et al. Private information leakage from single-cell count matrices. Cell 187, 6537–6549 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Öztürk, H. et al. Towards useful and private synthetic omics: community benchmarking of generative models for transcriptomics data. Preprint at bioRxiv https://doi.org/10.64898/2026.03.02.707794 (2026).Walsh, C. G., Sharman, K. & Hripcsak, G. Beyond discrimination: a comparison of calibration methods and clinical usefulness of predictive models of readmission risk. J. Biomed. Inform. 76, 9–18 (2017).Article  PubMed  PubMed Central  Google Scholar Xian, R. P. et al. Robustness tests for biomedical foundation models should tailor to specifications. NPJ Digit. Med. 8, 557 (2025).Article  PubMed  PubMed Central  Google Scholar Becker, J., Rush, N., Barnes, E. & Rein, D. Measuring the impact of early-2025 AI on experienced open-source developer productivity. Preprint at https://doi.org/10.48550/arXiv.2507.09089 (2025).Challapally, A., Pease, C., Raskar, R. & Chari, P. State of AI in business 2025: the GenAI divide. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf (2025).Mahmood, F. A benchmarking crisis in biomedical machine learning. Nat. Med. 31, 1060–1060 (2025).Article  CAS  PubMed  Google Scholar Wang, F. The crisis of biomedical foundation models. J. Biomed. Inform. 171, 104917 (2025).Article  PubMed  Google Scholar Lobentanzer, S., Rodriguez-Mier, P., Bauer, S. & Saez-Rodriguez, J. Molecular causality in the advent of foundation models. Mol. Syst. Biol. 20, 848–858 (2024).Article  PubMed  PubMed Central  Google Scholar Borges, J. L. A Universal History of Infamy (Penguin, 1975).Mermin, N. D. Could Feynman have said this? Phys. Today 57, 10–11 (2004).Google Scholar Rodríguez-Moris, G. & Docobo, J. A. The discovery of Neptune revisited. Preprint at https://doi.org/10.48550/arXiv.2405.06310 (2024).Kuhn, T. S. The Structure of Scientific Revolutions 3rd edn (University of Chicago Press, 2009).Kazmierczak, R., Berthier, E., Frehse, G. & Franchi, G. Explainability and vision foundation models: a survey. Inf. Fusion 122, 103184 https://doi.org/10.1016/j.inffus.2025.103184 (2025).