I've been working on clinical speaker attribution at Omi and wanted to compare the current diarization models on the same audio. I used 15 mock doctor–patient consultations from PriMock57, about 2.4 hours. Full recordings, automatic speaker counts, without telling the models there are two people. Batch Diarization error rate (DER), with ±250 ms boundary tolerance. Lower is better. Model DER Median processing time Pyannote Precision-3 2.891% 18.9 s / recording (API) Nemotron 3 4.803% 0.688 s / recording Pyannote Community-1 6.620% 18.691 s / recording Sortformer v1 6.778% 3.869 s / recording Sortformer v2.1 7.974% 1.077 s / recording VibeVoice-ASR 8.233% 123 s / recording Meta Muse Voice Transcribe † 13.042% 92 s / request (API) Local models ran on one NVIDIA L4. API times include round-trip overhead; VibeVoice-ASR also performs transcription. † Muse used 20 separate clips because of its 10-minute request limit, so its result isn't a whole-recording comparison. Pyannote Precision-3 had the lowest error. Nemotron came next and was the fastest local model. Streaming Model DER Pyannote live API 3.959% Nemotron 3 † 4.971% Sortformer v2.1 † 6.958% VibeVoice 1.5B 17.210% VibeVoice 7B 18.032% † Native streaming presets evaluated through unpaced, completed-file replay. The other rows use paced, delivered speaker outputs. These scores don't establish live latency. I didn't evaluate Muse for streaming. Same weights, different runtime I also tried optimizing Nemotron and Community-1 with our proprietary runtime, without changing the weights: Nemotron: 4.803% → 3.174% DER. 34% lower error, 2.13× faster. Community-1: 6.620% → 5.435% DER. 18% lower error, 24× faster. With zero boundary tolerance, Nemotron's runtime result gets slightly worse: 12.720% → 13.203%. Both scores are published. It's a small set with VAD-refined references, and we developed the runtime settings on it. Audio, references, scorer, saved outputs and NVIDIA baseline runners are public. Our runtime code stays private, but its outputs are included for rescoring. Repo and write-up in the comments. Any other diarization models worth adding?   submitted by   /u/MajesticAd2862 [link]   [comments]