I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp

Wait 5 sec.

Hi All: I have an Asus ESC 4000 G3 with 128 GB DDR4 RAM - I tried putting in my V620s but couldn’t put more than 2. Sadly, pivoted to T4s, these are 72 watt passive cards and I thought I could use them like how I use the Mi50 32GB - but I was very wrong. It has no support from Nvidia when it comes to latest fp8 emulation, reason being that it lacks resources. I am not sure if special vLLM repos exist, but I am trying to serve Qwen3.5-9B-AWQ-INT8 with vLLM or something faster than llama.cpp. I currently have 8 separate instances and they’re serving Qwen3.5-9B-Q5_K_XL at 87k per slot and there are 44 slots. So, my agentic (non-coding) harness works, its able to fetch emails, summarize documents, and do a lot for me and my team, but the concern is latency. Each slot continuously generates 25-40 tps (depending on the question), and essentially prefill is at 1400 tps, so I am not sure what my bottleneck is. VLLM on my Mi50 32GB and AWQ-INT8 (cyankiwi’s model) is very fast. I am wondering, does anyone have any pointers? I am really not in the position to spend money, exhausted everything for the next 6 months to a year already. I would greatly appreciate your help.   submitted by   /u/exaknight21 [link]   [comments]