llm-tune: getting local models that actually perform

Wait 5 sec.

Made a little agent skill called llm-tune to help find the best settings for running local LLMs on your hardware. Still working on it, but I'm looking to test it across more GPU setups. If you try it out, feedback and benchmark results are welcome. Supported: Architectures: Dense, MoE, hybrid MoE/Mamba GPUs: NVIDIA, AMD, Intel Arc Apple Silicon: M-series Macs via MLX Engines: llama.cpp, Ollama, vLLM Tuning: Quantization, context/KV cache, GPU offloading, MTP, sampling, reasoning, and agent harness settings Benchmarks: VRAM/RAM usage, tokens/sec, context recall, and output quality Currently measured on an RTX 4090; other hardware and backends are documented but need more real-world testing.   submitted by   /u/Distinct-Pie2389 [link]   [comments]