Radford, A. et al. Learning transferable visual models from natural language supervision. In Proc. 38th International Conference on Machine Learning 8748–8763 (2021). This paper established contrastive image-text pretraining (CLIP) as a foundation for modern vision-language models.Jiang, L. Y. et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357–362 (2023). This study demonstrated that large-scale health system data can support broadly useful clinical prediction models.Article CAS PubMed PubMed Central Google Scholar Assran, M. et al. Self-supervised learning from images with a joint-embedding predictive architecture. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 15619–15629 (2023). This paper introduced I-JEPA, the self-supervised representation learning method adapted here for volumetric medical imaging.Lyu, Y. et al. Learning neuroimaging models from health system-scale data. Nat. Biomed. Eng. https://doi.org/10.1038/s41551-025-01608-0 (2026). This paper showed strong performance with neuroimaging learning from MRI–report pairs at health system scale.Article PubMed Google Scholar Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023). This perspective outlines the rationale for generalist medical AI models that learn across heterogeneous clinical data modalities and support multiple downstream tasks.Article CAS PubMed Google Scholar Download references