Objective Although numerous clinical prediction models (CPMs) have been developed to identify pregnancies at risk of serious adverse outcomes, most lack robust external validation, which restricts their reliability in clinical practice. This study aims to systematically review and externally validate existing clinical prediction models for Gestational Diabetes Mellitus (GDM), Pre-eclampsia (PE), Stillbirth, and Small-for-Gestational-Age (SGA) focusing on models for use in pre-conception or early pregnancy. The models were evaluated using a large, representative primary care cohort to assess their performance within a UK population. Method and Analysis A two-stage systematic review identified CPMs for GDM, PE, Stillbirth, and SGA. Eligible models used routinely collected maternal characteristics, fully reported the model equation, and included predictors commonly measured during the preconception or early pregnancy period. Validation was performed using the UK Clinical Practice Research Datalink (CPRD Aurum), including over 1.58 million pregnancies for women aged 14 - 49 between 2000 to 2020. Model performance was assessed through discrimination (C-statistic), calibration-in-the-large (CITL), calibration slope and calibration plot, with pooled estimates across 20 imputations. Results Of 30 studies included, 48 models were identified, comprising 47 binary outcome models and one continuous outcome model, and 23 models used UK cohorts. Models were developed in cohorts comprising 101 to 113,415 participants (median N=5,013). Whilst 29 models have been externally validated, only three have been validated in cohorts including UK population satisfying the requisite sample size. Across all outcomes, most models demonstrated limited generalisability in the UK population. The C-statistic for GDM ranged from 0.29 to 0.77, for PE from 0.42 to 0.71, whilst models for stillbirth and SGA showed weaker discrimination from 0.54 to 0.56 and 0.58 to 0.62, respectively. Models across all outcomes exhibited substantial miscalibration, with calibration-in-the-large ranging from -5.66 to 1.35, and calibration slope from -0.65 to 17.61. Conclusion Existing CPMs for pre-conception and early prediction of adverse pregnancy outcomes showed poor discrimination and frequent miscalibration in a large UK cohort. Whilst some models for SGA had excellent performance, most existing models for GDM, PE, and stillbirth demonstrated suboptimal generalisability, likely a result of differences in population characteristics and outcome prevalence. Therefore, most CPMs evaluated are not suitable for direct implementation in UK clinical practice and require further calibration, updating, even re-development using large, representative data before achieving clinical utility.