Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

Wait 5 sec.

Following on from club-5060ti and club-rdna16, I’ve put together Infermeld: a small, open-source Linux companion kit for running one GGUF across an AMD GPU and an NVIDIA GPU, powered by llama.cpp. I’m the maintainer. This is an experimental v0.1.0 release, and I’m looking for people with other mixed GPU combinations to help reproduce the setup and find the rough edges. The idea is practical: if you already have cards from both vendors, can you put them to work together without buying a matching pair? What Infermeld adds The inference engine is llama.cpp. Infermeld isn’t a new backend, and I’m not claiming to have invented mixed-GPU inference. It packages the supporting pieces around that setup: Explicit AMD/Vulkan + NVIDIA/CUDA device selection and runtime preflight. Reproducible build instructions and inspectable launch arguments. A read-only thermal guard, with shutdown limited to the server process it started. Documentation and a results site that keep configurations, failures and limitations visible. The release is source-only. You build the documented llama.cpp revision separately and supply your own model weights. It’s intended for people comfortable with an experimental Linux setup, not as a one-click installer. Current tested setup Component Tested configuration AMD GPU RX 6900 XT, 16GB NVIDIA GPU RTX 3080, 10GB Model Qwen3.6-35B-A3B, UD-Q4_K_M GGUF Backends Vulkan + CUDA Split mode Layer Context reservation 8,192 tokens The acceptance checks include loading and short completions with MTP off and on. That’s a narrow result on one hardware pair, not broad compatibility testing. An 8K context reservation is not the same as testing a filled 8K prompt, and a short successful response is not a sustained performance benchmark. Important limitations Sustained Q4 throughput and full-length high-context results are not yet qualified. Historical measurements are labelled with their original configurations. They should not be read as performance numbers for the current Q4 setup. There’s no promise that combining cards is faster than using one. Adding the advertised VRAM capacities does not guarantee that all of it is usable for the model and its runtime allocations. I’d rather make those boundaries clear than present a successful load as a complete benchmark. Looking for other AMD/NVIDIA combinations Successful runs and failures are both useful. If you try it, please include: Both GPU models and their VRAM sizes. OS, driver versions and llama.cpp revision. Model and quantization. Launch settings, including the split and context reservation. How far it got: preflight, loading, first completion or a longer workload. There’s a hardware/result issue form in the repository. Please sanitize paths and keep credentials and private logs out of reports. Repository and setup instructions: https://github.com/5p00kyy/infermeld Results and evidence: https://5p00kyy.github.io/infermeld/ Anyone already using an AMD/NVIDIA pair for local inference? I’d be interested in what works for you, and where this setup breaks on different hardware.   submitted by   /u/do_u_think_im_spooky [link]   [comments]