I built an open-source meeting assistant that lets me query a 2-hour debate locally on an 8GB GPUTexto:

Wait 5 sec.

I've been working on Meet2Notes, a local meeting assistant, and I’m pretty happy with this latest test. I fed it a recording of a two-hour political debate: seven party leaders, one moderator, and a lot of different arguments. After transcription and loading the transcript into the local LLM, I could ask questions about what each person said. In the demo, follow-up answers start appearing in around 1–3 seconds, running Bonsai 27B 1-bit on an RTX 3070 with 8GB VRAM and 32GB system RAM. No cloud API was used for this workflow. The important detail: that’s time to first text with the context already loaded, not the time to process the recording. The initial cold request took about 76 seconds, after transcription. I've spent the last few days testing transcription and speaker diarization models, adjusting memory usage, and getting the model to reuse its cached context instead of processing the whole transcript again for every question. Meet2Notes can record or import meetings, transcribe them, separate speakers, generate notes, and let you ask follow-up questions. I’ve just released v0.9.0, including faster saved-voice matching and a smaller Bonsai 8B option. It’s free and open source. There’s still plenty to improve, especially across different hardware and recordings, but seeing this work locally has been a satisfying milestone. GitHub Full demo If you work with long meetings, lectures, or interviews, what would you test first? I’d especially appreciate feedback on setup and answer accuracy.   submitted by   /u/VERSATILCORDOBA [link]   [comments]