Medical AI’s accuracy gains are outrunning evidence of better patient outcomes

Wait 5 sec.

A plethora of studies in 2026 is showing a persistent gap in medical AI: the systems can impact doctors’ choices and achieve impressive diagnostic results without demonstrating that patients actually become healthier. This is important for hospitals purchasing these systems, companies trying to showcase their usefulness, and authorities determining which evidence should be considered once AI is adopted in doctors’ clinics.The proof problem regulators are starting to scrutinizeThe disparity between the remarkable performance of AI and proven efficacy for patients appears to be becoming too big to overlook. The Financial Times summed up this situation in its headline:“Medical AI has a proof problem.”This issue has now moved from the realm of academic discussions. The FDA is asking many of the same questions as it looks at how to assess these AI-enabled medical devices before and after they are put to use. The August 18 discussion paper, open for public input until 19 October, covers such issues as risk assessment, premarket assessment, and postmarket monitoring.The FDA is careful to make clear that this paper is purely a discussion paper and not a draft guideline, a proposed change in policy or any indication of future regulatory expectations.A previous FDA’s inquiry regarding real-world performance has also brought into consideration the issue of model drift and the limitations of using static metrics. This raises a larger question that cannot be ignored anymore: how do we measure safety and efficiency once an AI solution has moved out of the laboratory?In Kenya, AI support did not lower treatment failureIn a study published in the Nature Medicine journal, which was released on June 26, 103 clinical officers at 16 facilities of Penda Health in Nairobi and Kiambu counties treated patients with or without the help of a large language model (LLM) for decision-making. Out of the 9,691 patients registered, 9,347 patients were considered for the primary analysis.The rate of treatment failures after 14 days was 2.2% in the AI group and 2% in the control group so the adjusted odds ratio was 0.77, although the difference is not statistically significant (P=0.13). No serious adverse events were associated with the intervention.The group using AI did have better documentation and more appropriate diagnoses and treatment plans. The significance of the result derives from the fact that better process measures did not result in any improvement in the outcome for patients.The trend is similar to the RAPIDx AI trial of 2024. In the primary analysis of 3,029 patients, the composite six-month outcome of cardiovascular death, myocardial infarction, or unplanned cardiovascular readmission was 26% using AI and 26.4% using standard care. However, within the non-type 1 MI group, invasive coronary angiography was performed with 47% lower frequency with AI-supported care.The scores keep climbing; the real-world proof lagsA multi-country randomized trial under Nicholas Rounding’s guidance found that GPT-4o improved 249 doctors’ clinical vignette performances by 18% in Kenya, 10.7% in Indonesia, and 7.2% in the Netherlands (P