by Jack Kennedy, William Ferguson, Owen Jones, Steven Riley, Thomas Ward, Maria L. Tang, Jonathon MellorBackground Epidemic forecasting research often assesses ensembles and their component models using probabilistic scoring rules. Quantifying how individual models affect ensemble performance is challenging, particularly across multiple targets and spatial scales. Methods We present Winter 2024–25 forecasts of Influenza and COVID-19 hospital admissions in England and conduct a retrospective simulation using the operational component models. Forecasts were scored using the per capita weighted interval score (pcWIS) for counts and the ranked probability score (RPS) for ordinal trend direction. We compared retrospective forecasts, used generalised additive models (GAMs) to estimate the expected change in score from the inclusion of a model in a sub-ensemble (an ensemble formed from a subset of available models), and used Pareto analysis to understand which sub-ensembles were Pareto-optimal across scoring rules. Results Nationally, there was a 47% improvement in Influenza pcWIS versus sub-ensembles. However, Influenza operational ensembles were on average 22% worse than sub-ensembles, when measured by RPS. For COVID-19, operational ensembles were 43% and 280% worse on average, than retrospective sub-ensembles by pcWIS and RPS, respectively. However, COVID-19 operational ensembles were on average 2% (pcWIS) and 13% (RPS) better than individual operational models. For influenza, operational ensembles were, on average, 58% (pcWIS) and 41% (RPS) better than individual models. The sub-ensemble simulation showed how individual models influenced the ensemble scores during different epidemic phases. The Pareto analysis demonstrated that there can be a trade-off between relative direction and absolute count score optimisation.