GPU - Vulkan llama.cpp benchmarks sorted by price to performance

Wait 5 sec.

This table to help anyone looking to build a budget Data Center homelab. I copied the bulk of value based, mid level, decent speed results GPUs and feed it to AI or SI and here are the recommended results. Data taken from Llama.cpp discussion thread: Performance of llama.cpp with Vulkan #10879 There are 76 different GPU models listed in the benchmark. "Testing the 'Llama 2 7B model' and use Q4_0 as it's simple to compute and small enough to fit on a 4GB GPU" Based on the specific llama-bench baseline data provided, running local LLM inference via the Vulkan backends shifts the value hierarchy drastically. Modern mid-range consumer cards are severely bottlenecked by narrow bus widths (128-bit or 192-bit) during decoding (tg128), whereas enterprise components and older massive-bus flagships dominate performance-to-cost value. By analyzing the current 2026 secondary market pricing (collating active trends across secondary platforms like eBay and specialized tech hardware communities) against your baseline metrics, here is the performance-to-cost value ranking. The cost-to-performance efficiency formula balances the entry price against prefill speeds (pp512), decoding throughput (tg128), and total accessible VRAM. Top 20 GPU Performance-to-Cost Ranking (Used Market) Rank GPU Model Est. Used Price pp512 (t/s) tg128 (t/s) VRAM Capacity Performance-to-Cost Architecture Profile 1 Nvidia P102-100 ~$40 - $50 ~510 ~62.8 10 GB Absolute Value King: Stripped mining card with a 320-bit bus. Yields ~1.3 tokens/sec per dollar spent on decode cycles. 2 AMD Instinct MI50 ~$110 - $130 ~1,119 ~108.5 16 GB tg128 Efficiency King: Full 1,024 GB/s HBM2 bandwidth. Best cost-per-token decode engine on the secondhand market. 3 AMD Radeon VII ~$140 - $160 ~1,059 ~101.1 16 GB Same elite HBM2 memory substrate as the MI50 but packaged with consumer display outputs. 4 Nvidia GTX 1080 Ti ~$110 - $130 ~585 ~67.7 11 GB Legacy consumer warrior. Its wide 352-bit bus regularly out-decodes modern architecture under $300. 5 Nvidia Tesla P100 ~$90 - $110 ~678 ~63.1 16 GB Budget HBM2 alternative. Slower core processing bounds its prefill, but decode values are incredibly high. 6 AMD Radeon RX 6800 ~$220 - $240 ~1,593 ~101.4 16 GB Exceptional balance. Clean driver architecture yields massive decode velocity relative to modern hardware tiers. 7 AMD Radeon RX 7900 GRE ~$400 - $430 ~2,336 ~116.1 16 GB Modern value standout. RDNA3 architecture scales beautifully on compute tasks with excellent memory throughput. 8 Nvidia RTX 3060 (12GB) ~$180 - $200 ~1,815 ~75.9 12 GB The entry-level standard for consumer setups. Ample VRAM budget for small models at a highly accessible price tier. 9 Nvidia Tesla V100 (16GB) ~$180 - $220 ~1,391 ~129.5 16 GB Combined Enterprise Pick: Volta core structure provides blisteringly reliable generation and prefill baselines. 10 Nvidia RTX 2080 Ti ~$200 - $230 ~1,888 ~97.5 11 GB Highly efficient Turing flagship layout. Out-paces newer equivalents due to an aggressive 352-bit bus framework. 11 AMD Radeon RX 7800 XT ~$350 - $380 ~2,017 ~118.2 16 GB Clean, highly competitive RDNA3 compute engine displaying great out-of-the-box Vulkan metrics. 12 AMD Radeon RX 7900 XT ~$500 - $550 ~2,941 ~123.1 20 GB Massive 20GB framework buffer size. Excellent performance scale, though commands a higher price footprint. 13 Nvidia Tesla P40 ~$120 - $140 ~488 ~59.3 24 GB The cheapest entry to 24GB allocation. Let down by poor FP16 computing speeds, keeping context loading sluggish. 14 Nvidia RTX 4070 Super ~$480 - $520 ~4,608 ~108.7 12 GB Blistering prefill speed bounds. Highly performant cores make up for the standard 192-bit bus structure. 15 Intel Arc A750 ~$90 - $110 ~1,075 ~42.6 8 GB Phenomenal raw bandwidth per dollar, but tightly restricted by a fixed 8GB VRAM ceiling. 16 Nvidia RTX 4070 Ti Super ~$680 - $730 ~6,099 ~129.4 16 GB Outstanding raw throughput benchmarks, but hits a higher tier of up-front investment cost. 17 Nvidia RTX 5060 Ti ~$420 - $460 ~3,460 ~93.5 12 GB / 16 GB Blackwell mid-tier layout. Offers highly robust processing bounds, though carries a modern market premium. 18 AMD Radeon RX 580 ~$40 - $50 ~258 ~39.3 8 GB Dirt cheap entry floor. Delivers text processing capability at the lowest possible cost parameter. 19 Nvidia P104-100 ~$30 - $40 ~311 ~46.1 8 GB Low-profile budget node. Useful for multi-card distributed matrices where base components must be inexpensive. 20 AMD Radeon RX 9070 ~$550 - $600 ~3,164 ~119.7 16 GB Next-gen RDNA4 architecture architecture layout. High performance density but subject to lower hardware-to-cost scaling. Key Strategic Takeaways from Vulkan results Lowest Cost for Token Generation (tg128): The AMD Instinct MI50 and P102-100 completely distort the curve. The MI50 nets you over 100 t/s on a Llama-7B architecture for roughly $120, a metric that consumer desktop tiers require twice the budget to replicate. Lowest Cost for Prompt Processing (pp512): Modern architectures rule prefill metrics due to hardware tensor capabilities. If prompt processing latency is your critical bottleneck, look at the RTX 4070 Super or RTX 5060 Ti, which punch far above their weight class on ingest speeds. Best Combined Balancer: The GTX 1080 Ti and AMD Radeon RX 6800 hit the absolute "sweet spot" for standard desktop nodes. They avoid the strict cooling modifications or specialized software handling required by headless data center units (like the Tesla series) while maximizing bandwidth-to-dollar efficiency. Chart and summary provided by Gemini and myself. I currently own RX 7900 GRE, MI50, P102-100, GTX-1080Ti, GTX 1070, RX 480/580. Here is the breakdown of the cost-per-token-per-second (\(\div \text{t/s}\)) for each metric across the top 20 GPUs. Lower cost values ($/t/s) mean you get more performance out of every dollar spent. Combined throughput represents a balanced arithmetic baseline of both prefill and generation. GPU Model Est. Used Price pp512 Cost per t/s tg128 Cost per t/s Combined Cost per t/s Nvidia P102-100 $45 $0.0881 $0.7162 $0.1569 Nvidia P104-100 $35 $0.1122 $0.7579 $0.1955 AMD Instinct MI50 $120 $0.1072 $1.1059 $0.1954 AMD Radeon RX 580 $45 $0.1744 $1.1445 $0.3027 Intel Arc A750 $100 $0.0929 $2.3441 $0.1788 Nvidia RTX 3060 $190 $0.1046 $2.5020 $0.2009 Nvidia RTX 4070 Super $500 $0.1085 $4.5981 $0.2120 Nvidia RTX 2080 Ti $215 $0.1139 $2.2033 $0.2165 Nvidia RTX 4070 Ti Super $705 $0.1156 $5.4461 $0.2264 Nvidia RTX 5060 Ti $440 $0.1271 $4.7054 $0.2476 AMD Radeon VII $150 $0.1416 $1.4824 $0.2585 Nvidia Tesla V100 $200 $0.1437 $1.5434 $0.2630 Nvidia Tesla P100 $100 $0.1475 $1.5833 $0.2698 AMD Radeon RX 6800 $230 $0.1443 $2.2667 $0.2713 AMD Radeon RX 7900 GRE $415 $0.1776 $3.5742 $0.3384 AMD Radeon RX 7800 XT $365 $0.1809 $3.0862 $0.3418 AMD Radeon RX 7900 XT $525 $0.1785 $4.2621 $0.3426 AMD Radeon RX 9070 $575 $0.1817 $4.8033 $0.3502 Nvidia GTX 1080 Ti $120 $0.2049 $1.7712 $0.3674 Nvidia Tesla P40 $130 $0.2664 $2.1900 $0.4750 The 10 worst GPUs based on performance-to-cost value are ranked below using the provided benchmark dataset and current secondhand market value trends. These values represent the highest cost per token per second ($/t/s). A higher number means you are paying significantly more money for every unit of inference speed generated. Rank GPU Model Est. Used Price pp512 Cost per t/s tg128 Cost per t/s Combined Cost per t/s Primary Bottleneck Profile 1 Nvidia Tesla M40 $60 $0.6488 $1.5248 $0.9103 Worst Overall Value: Outdated Maxwell architecture yields critically low processing throughput across both prefill and generation. 2 Nvidia Titan V $350 $0.4395 $3.3314 $0.7766 Premium Collector Tax: Despite HBM2 memory, a high up-front market premium makes its performance-to-dollar ratio poor. 3 AMD Radeon Instinct MI60 $150 $0.4062 $1.9191 $0.6705 Severely low prefill scaling limits its deployment utility relative to the much cheaper MI50 framework. 4 AMD Radeon RX 7600 XT $260 $0.3092 $4.9038 $0.5817 Extreme Decode Bottleneck: A very narrow 128-bit bus forces an incredibly inefficient $4.90 per token/sec on decode loops. 5 AMD Radeon RX 6600 XT $160 $0.2784 $2.9674 $0.5091 Limited by entry-tier bandwidth configurations that fail to translate into meaningful compute value. 6 Nvidia Tesla P40 $130 $0.2664 $2.1900 $0.4750 While popular for cheap 24GB capacity, missing native FP16 compute hardware tanks its relative speed value. 7 AMD Radeon RX 5700 XT $130 $0.2454 $1.8379 $0.4330 Older RDNA1 compute layers drop performance significantly compared to modern secondhand equivalents under $150. 8 AMD Radeon RX 6900 XT $400 $0.2104 $3.7037 $0.3982 Commands a high premium on the used market but struggles to scale its text generation speeds efficiently. 9 AMD Radeon RX 6750 XT $220 $0.2114 $2.6836 $0.3920 Tightly squeezed by low raw compute density relative to its market price window. 10 Nvidia GTX 1070 $70 $0.2177 $1.6876 $0.3856 The basement floor of the Pascal generation. Replaced entirely by the vastly superior cost-to-performance curve of the P102-100. Tesla P40 made both charts.   submitted by   /u/tabletuser_blogspot [link]   [comments]