Liquid just dropped d1-omni-600M and it's a decision model! I ported it to runntime, the WebGPU inference library I'm currently working on. It's plain TypeScript on top of TypeGPU with no WASM and virtually no export step. The model is written directly from our core ops (matmul, attention, norms, a few elementwise bits), and the weights load straight from the HF safetensors. The demo in the video is a fake comment feed being moderated live. Each comment gets 4 questions: toxic? spam? asking something? overall tone? Toxic and spam ones get removed. ~180 ms per comment for all 4 questions, ~45 ms per question - the performance will most likely be way better once we spend some time tuning the engine for this model I also tried making it play snake, but It did not go well lol. The d1 port isn't in the npm release yet. The rest of runntime is (detection, segmentation, speech-to-text, embeddings and more). Docs and live demos: https://docs.swmansion.com/runntime Happy to answer questions about the port or WebGPU stuff in general.   submitted by   /u/FinancialAd1961 [link]   [comments]