We've been running our own LLM for the team for about a week, and I wanted to share how it turned out, since "can one machine actually serve a whole team?" comes up here a lot. Short answer: yes, if you pick the right model. Hardware and model One NVIDIA DGX Spark (128GB unified memory) Ornith-1.5-35B-A3B, an MoE with 35B total and 3B active parameters. The 3B active is the whole trick: you get decent quality while the box only does about 3B worth of compute per token. Official 4-bit NVFP4 checkpoint, about 22 GiB, so no quantizing it yourself vLLM 0.24 with multi-token prediction (speculative decoding with no separate draft model) What it does One person gets 79 tok/s, which feels instant in a chat UI 20 people at once get about 25 tok/s each, 490 tok/s total Every request can use the full 262K context Vision and tool calling both work Reasoning scores: MATH-500 97%, GSM8K 98%, MMLU-Pro 81%. That's not frontier level, but for daily coding help, docs, and Q&A, nobody on the team has complained. How people use it OpenCode in the terminal for coding Open WebUI for chat and PDFs. PDF parsing and embeddings run locally too. Nothing goes to an outside API. For us that was the whole point, since a lot of our code and documents can't leave the building. The part nobody talks about: multi-user Running a model for yourself is easy. Running it for a team means you need to answer "who's using it, how much, and is someone hammering it?" So vLLM isn't exposed at all. Everything goes through a small gateway that: Gives each person and device their own API key Counts input and output tokens per user live Pushes the counts to Grafana dashboards Alerts if anything tries to talk to vLLM directly It's a small Python service, and honestly it's the piece that made this feel like real infrastructure instead of a side project. Gotchas if you try something similar On unified memory, don't get greedy with GPU memory utilization. I had to cap it at 0.70, or concurrency fell apart under load. More concurrent slots isn't always better. Past 20 users, throughput dropped because requests started queuing. Open WebUI stores the default model in its own database and ignores the env var after first setup. I lost an hour to that. Everything's on GitHub: the setup scripts, gateway, Grafana dashboards, and benchmark scripts with raw results. https://github.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup If you're thinking about self-hosting for a small team, happy to answer questions. I'm also curious what others use for per-user tracking. I looked at LiteLLM but wanted something tiny that I fully understood. https://preview.redd.it/co3i4tigirth1.png?width=2160&format=png&auto=webp&s=d61da076bcd37e5f91dff7f298de0d09e1dd6638   submitted by   /u/Commercial_Designer5 [link]   [comments]