A Framework to Monitor Editing of Artificial Intelligence-Generated Medical Documentation

Wait 5 sec.

Objective: To develop and evaluate a framework for characterizing clinician editing of Artificial Intelligence (AI)-generated[KS1.1] documentation and assess its feasibility for health system-level monitoring of AI scribes. Materials and Methods: We analyzed outpatient encounters in which an AI scribe was used at a single academic health system, examining the History of Present Illness (HPI) and Assessment and Plan (A&P) sections of notes. We characterized edits on three dimensions[KS2.1][AG2.2]: lexical edit intensity based on Levenshtein distance, embedding edit intensity BERT[KS3.1]Score, and clinical edit intensity based on removed and added UMLS[KS4.1][AG4.2] concepts.[KS5.1][AG5.2] We analyzed the Positive Predictive Value (PPV) of clinical edit intensity as a measure of clinically meaningful editing using clinicians as the gold standard, examined correlations among dimensions, and designed exponentially weighted moving-average control charts to monitor longitudinal changes in clinician editing behavior.[KS6.1][AG6.2] Results: 268,379 encounters were included (267,654 HPI, 267,594 A&P). Clinical edit intensity [≥]1 had an 88.9% PPV for clinically meaningful editing. Lexical and embedding edit intensity were highly correlated (Spearman {rho} 0.95), while clinical edit intensity was less strongly correlated with both ({rho} 0.77-0.81). Longitudinal monitoring detected changes coinciding with system-wide rollout.[KS7.1][AG7.2] Discussion:[KS8.1][AG8.2] Clinicians edited A&Ps more heavily than HPIs, potentially reflecting greater attention to content involving clinical decision-making. Over one-third of sections involved clinical concept changes, and clinical edit intensity identified clinically meaningful edits while providing information complementary to lexical editing measures. Conclusion: Clinician editing can be characterized at scale using complementary editing dimensions, providing a scalable signal for post-deployment surveillance of the human-AI documentation process.