How Predictable Was the 2026 Bundibugyo Virus Disease Outbreak? A Rolling-Origin Evaluation of Short-Term Forecast Models, an Empirically Recalibrated Bayesian Model, and a Data-Driven Baseline-Bayesian Ensemble, Using Daily Surveillance Data

Wait 5 sec.

Background The 2026 Bundibugyo virus disease outbreak in the Democratic Republic of the Congo became the largest recorded epidemic caused by Bundibugyo virus and showed unusually rapid early growth. Initial projections warned that it could become one of the largest Ebola-family outbreaks on record, but these were based on assumed mortality totals and intervention scenarios rather than repeated comparison with subsequently observed surveillance data. We assessed how accurately the outbreak could have been forecast in real time from routinely published national situation reports, whether a simple ensemble improved on its component models, and how forecasts should inform operational capacity planning. Methods We conducted a retrospective pseudo-prospective rolling-origin study using daily cumulative confirmed cases and deaths reported from 14 May to 20 July 2026 (68 calendar days; 61 numeric reports). Each date with a newly reported national total became an eligible forecast origin once seven numeric observations were available, yielding 55 origins. At each origin, all later data were withheld. Two prespecified models were refitted using only information then available: a recent seven-increment baseline and a Bayesian negative-binomial surveillance-maturity model fitted by Markov Chain Monte Carlo. A Gompertz model was fitted identically as a benchmark. Forecasts were produced for 7, 14, and 21 days and evaluated only when an observation existed on the exact target date. Bayesian forecasts were prospectively recalibrated using only previously realised errors from forecast-maturity-qualified origins. We also evaluated a baseline-Bayesian ensemble, with horizon- and outcome-specific weights selected prospectively by an expanding-origin procedure using an asymmetric operational loss function. Findings A statistically and externally corroborated surveillance-maturity discontinuity occurred on 28 May 2026 (robust z score 13{middle dot}0). The Gompertz model had the poorest 80% predictive-interval coverage at every horizon and for both outcomes (25{middle dot}9-34{middle dot}1% for cases; 0{middle dot}0-23{middle dot}5% for deaths) and was excluded. Recalibration corrected consistent Bayesian underprediction and was applied at 24 of 55 origins for 21-day forecasts. It substantially improved case coverage (7-day, 60{middle dot}0% to 80{middle dot}0%; 14-day, 41{middle dot}0% to 76{middle dot}9%) and improved death forecasting in both accuracy and calibration (7-day median absolute percentage error, 10{middle dot}3% to 7{middle dot}7%; 80% coverage, 33{middle dot}3% to 95{middle dot}6%). Ensemble weights differed by outcome: case forecasts were baseline-heavy (w{approx}0{middle dot}8-0{middle dot}9), whereas death forecasts were near parity (w{approx}0{middle dot}5-0{middle dot}6). For deaths, the ensemble improved mean weighted interval score over both component models at every horizon (7-day: baseline 24{middle dot}8, Bayesian 20{middle dot}5, ensemble 17{middle dot}9). Interpretation Routine situation-report data supported useful short-term forecasting, but no single model was best on every criterion. Three forms of maturity shaped forecast reliability: epidemic, surveillance, and forecast maturity. Forecast maturity was the most robust and model-independent finding and supports 21 days, rather than 28 days, as the longest routine operational horizon. The study provides a reproducible and adjustable framework for combining simple and complex models and communicating uncertainty to operational decision-makers.