Abstract Objectives To determine whether single-layer evaluation characterises clinical risk in the final note of a multilingual ambient AI scribe, and where serious errors arise. Design Two-arm evaluation on one common clinical-risk scale, using a frozen, reference-aligned synthetic corpus. Setting Scripted consultations spanning 77 languages and five complexity levels, from simple general practice to expert multidisciplinary-team handover. Main outcome measures Per-session incidence of at least one serious (HIGH or CRITICAL) note-layer error in intrinsic (reference script to note, n=385) and end-to-end (automatic speech recognition (ASR) transcript to note, n=2,302) generation; three LLM raters classified discrepancies using a four-mode taxonomy. Results Intrinsic note generation had a 3.4% serious-error rate, with no detectable gradient across complexity (1% to 6%; Cochran-Armitage p=0.42) or language resource (high 2%, medium 5%, low 4%). End-to-end serious errors rose steeply with complexity (L1 1% to L4/L5 24%) and language scarcity (high 7%, medium 10%, low 15%; OR 1.54 per tier, 95% CI 1.15-2.07; p=0.004, GEE clustered on language). Of serious in-note errors, 87% were ASR-derived and 7% note-originated; amplification exceeded correction threefold (fate entropy 1.19 bits). Repeated generation showed moderate reproducibility (Fleiss kappa 0.42); word error rate explained 24% of cross-language variance versus ~1% for transcription risk density. Conclusions Low serious-error rates in individual layers did not preclude higher end-to-end note risk. Evaluation should therefore include the clinician-facing note, because component-level metrics cannot fully capture risk created or transformed at the transcription-to-generation interface.