Objectives Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exists and validation cannot mean accuracy. We tested whether such judges are a defensible instrument: reproducible, within the envelope of expert disagreement, and non-differential across arms. Methods Pre-registered, blinded validation study nested in a multi-country simulation of ambient AI documentation (English setting), reported per GRRAS. Ten external clinicians independently adjudicated a stratified sample of 434 pipeline flags, retained and screen-discarded, blinded to note authorship, identification source, the pipeline's verdict and severity tier. Agreement used Gwet's AC1; proportions carry Wilson intervals. Three propositions were pre-specified: envelope parity, non-differential behaviour across arms, and concordance on consensus cases. Results All ten reviewers completed: 565 adjudications across 434 items, 131 of them double-rated. Inter-clinician agreement on genuineness was fair (raw 59%, 95% CI 50 to 67; AC1 0.24), leaving no human consensus to serve as truth. Judge-clinician agreement was 64% (95% CI 60 to 68), overlapping that interval. Behaviour was near-symmetric on contrast-critical metrics: kept-precision 74% for AI against 81% for clinician notes, and severity signed gap +0.06 against -0.09 tiers. One sub-metric was asymmetric: removed-confirmed 56% against 42%, so the screen over-removes more on clinician notes, a direction conservative to the parent contrast. On 77 consensus items the pipeline concurred on 70% (95% CI 59 to 79). Latent-class triangulation placed the genuine-error rate among flagged candidates at 68% (94% credible interval 48 to 83). Conclusions The judges behave as a consistent, near-non-differential, clinician-equivalent instrument. This licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions. It is not a claim of accuracy, which moderate consensus concordance and fair reliability preclude, and the genuine-error rate is best reported as an interval.