Title: I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside. ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV. What's running (one Debian LXC, one setup script): - OpenAI-compatible router, one API key, swaps models on demand - qwen2.5-coder 1.5B/7B, qwen2.5-7B, ornith-1.5-9B, qwen3-30B-A3B, qwen2.5-VL-3B - Images: SD 3.5 Medium + SD 1.5 via stable-diffusion.cpp, Whisper medium - Also the inference backend for an agent client with MCP tools Measured, Ryzen 5 5600G / 32 GB RAM / RX 580 2048SP, temp 0, max_tokens 300: | Model | tok/s | warm | cold (after swap) | |---|---:|---:|---:| | qwen2.5-coder-1.5b | 99.2 | 3.1 s | 11.0 s | | qwen2.5-vl-3b | 66.4 | 4.4 s | 14.9 s | | qwen2.5-7b | 34.2 | 8.8 s | 24.3 s | | ornith-1.5-9b | 28.0 | 10.9 s | 26.3 s | | qwen3-30b-a3b (18.6 GB, from RAM) | 21.3 | 14.1 s | 55.0 s | Images: SD 1.5 = 27.5 s, SD 3.5 Medium = 53.3 s per 512x512. Two takeaways for 8 GB owners: - Don't force -ngl 99: on the MoE it cost 8.1 vs 22.3 tok/s (VRAM 99.4% full). Auto-fit layers won by nearly 3x. - Model swap = llama-server restart = 7.9-38.5 s. Keep one model per session. Full method and raw numbers in docs/BENCHMARKS.md, including the agent breakdown (first message ~161 s on a 20k-token prompt, ~25 s after caching). https://github.com/AvilaCarlosDev/polaris-local-ai (MIT)   submitted by   /u/Heavy-Level-5215 [link]   [comments]