Language models reflect clinical evidence but fail to adapt it to patients

Wait 5 sec.

Background: Clinical language models must use evidence to produce the number and action required by a particular case. Whether numerical knowledge reliably becomes a correct clinical response is unclear.`n`nMethods: NUMBERS evaluated 16 model configurations on 1,300 questions linked to public clinical evidence. Linked experiments tested prevalence updating, patient-specific estimates and clinical actions. The direct-action experiment compared cutoff recall plus action selection with a supplied complete rule across 50 rules, five models and 7,500 calls.`n`nResults: Among diagnostic estimates outside the source-result tolerance, 77.1% remained within the evidence's central 80% predictive range. Models updated prevalence-dependent quantities correctly in 78.5% of comparisons with inputs and a calculation request, versus 33.7% with clinical wording. Supplying inputs and requesting calculation raised patient-level near-target answers from 29.7% to 78.8% across 2,340 pairs. Models selected an incorrect action despite stating a cutoff that implied the correct action in 413 of 3,750 recall-arm calls (11.01%; 95% confidence interval, 8.93 to 13.17). Supplying the complete rule raised action accuracy from 85.63% to 99.41%, an improvement of 13.79 percentage points (95% confidence interval, 11.65 to 15.95).`n`nConclusions: Models often produced evidence-consistent numbers but failed to adapt them to a case or act consistently with their own stated cutoff. Explicit inputs, calculations and complete rules substantially improved performance in controlled prompts.`n`nFunding: National Academy of Medicine, Agreement No. 2026A008797.