Best STT / or literal transcription model? Does that exist?

Wait 5 sec.

I'm putting together a video editing mcp but the problem is, e.g with parakeet v3, is it's trying to be too smart. It seems to guess at what makes a sentence, removes um's and such. But we need those. What better models are there for this usecase? are there any models that potentially also do tagging, like, (laughs), so it gives more context for speaking cadence and such?   submitted by   /u/Zeeplankton [link]   [comments]