Background: Several breast cancer (BC) risk prediction models have been developed to provide personal risk assessments. Though individually validated, their performance has not been systematically evaluated across a wide range of populations or ages. Methods: We harmonized individual-level baseline questionnaire data and incident BC diagnoses from 21 cohorts from North America, Europe, and Australia participating in the Breast Cancer Risk Prediction Project. Five-year absolute risk of invasive BC was estimated for five established risk prediction models using classical risk factors only. Discrimination was evaluated by area under the curve (AUC). Calibration was assessed using average and risk-decile specific expected to observed (E/O) ratios. Performance metrics were meta-analyzed across cohorts and models. Metaregression tested associations between cohort characteristics and performance metrics. Results: This analysis included 1,595,977 women aged 20-75 years, enrolled in studies between 1976-2015, with 19,062 (1.2%) invasive BC cases ascertained within 5 years from exposure assessment. Age-adjusted AUCs were similar across models and cohorts (pooled AUCs by model: 0.57-0.58), while E/O ratios varied substantially (pooled E/O ratios by model: 0.83-1.25). Overestimation was common among predicted high-risk individuals (>3%). No appreciable differences in model performance by cohort age, birth year, race, and variable missingness emerged. Calibration improved after assigning race-specific incidence rates. Conclusion: Existing BC risk prediction models provided similar risk discrimination across multiple cohorts, although there was overestimation of risk for high-risk individuals. Performance variation across cohorts was not driven by specific characteristics, which supports development of a unified risk model for diverse populations that leverages appropriate incidence rates.