Offline reinforcement learning (RL) provides a promising framework for learning and evaluating treatment policies from logged clinical data, particularly in sequential decision-making settings where prospective exploration would be unsafe. In ICU sepsis management, however, it remains unclear whether offline RL policies retain stable behavior under increasingly severe out-of-distribution (OOD) patient cohorts. In this paper, we evaluate standard offline RL methods on three severity-enriched OOD test mixtures from the MIMIC-III benchmark dataset to determine whether offline policies retain a stable, actionsensitive decision-support signal. Under the shared learned-dynamics offpolicy evaluation (OPE) protocol, as the severe-OOD ratio increases from 25% to 75%, observed clinical survival declines from 67% to 49%, while the best offline method in each mixture receives model-predicted terminal survival values of 87%, 86%, and 85%, respectively. Because observed clinical survival and model-predicted terminal survival are different quantities, this contrast suggests a stable model-based decision-support signal under severity shift. We further present a secondary physiological stabilization analysis using an episode-level physiological stabilization score (EPSS), a heuristic summary of whether selected physiological variables move in favorable directions during follow-up. In this analysis, model-generated rollouts under offline policies receive higher EPSS values than matched logged clinical trajectories for several physiological components. Together, these results support learned-dynamics OPE as a useful severity-OOD stress test for offline RL policies in ICU sepsis, while leaving prospective and causal validation as necessary next steps.