Good day data nerds! I am trying to get more into running models, learning about agentic workflows, and creating my own tools. But as you do (right?) I had to fine tune my current setup. With the way the markets are right now, it only makes sense to make the most out of what I got. My day job, typically, is around data, numbers, and programming; only two of those I am good at, I'll let you guess which ones. So, yes, here come some data sheets and pretty graphs. Don't worry, you don't have to go through the data, but you can if you want. The graphs cover key metrics spat out by Ollama such as the tokens per second, duration, eval rates etc. I also added a cheeky "tokens/s/W" which, technically is not perfect since I do not measure wattage over time, but I did observe the watts during prompts, and I have a few things to mention about that later. Okay, let's start off with the specs because you probably think I am rocking the good stuff since I am so invested in this topic (haha) Lenovo M920Q 16GB DDR4 2660MHz Intel i5-8500T (6C/6T) Gigabyte Gaming OC 3070 8GB (Over OcuLink at Gen3x4 speeds) I am running OpenWebUI with Ollama in an LXC on my Proxmox server. This is one of my nodes and it is dedicated to my models. I have given it all of the cores, 14GB of RAM, 2GB swap. Nothing crazy to write home about, see? Okay, so, one of the things I wanted to know what, with the models that I run day to day, how effective are they at different GPU power limits. Man, if only I had known how much of a rabbit hole I would go down to do this (sorry wife). I only run 4 models, nothing too crazy, until you run 3 tests per model, for each power limit from 100 to 220 (156 runs in total), and each run of each model taking around 3 or 4 minutes since I have to unload the model each time to not have any prompt caching. Afterwards I would average the results and add that to the sheet. Really gave the fingers a workout since I now am a proud owner of a 60% keyboard for the first time and I no longer have a numpad... I'll remember that for next time. That being said, the switches are soooo creamy, a valiant tradeoff. So what models am I running? Glad you asked. The 3070 does limit me quite a bit, but with so many models available and so many smarter people than me who can quantize the models, I have found these models fit my needs. For the most part everything runs in the VRAM, except for 2, but those come with asterisks. gemma4:e4b qwen3.5 qwen3-vl* deepseek-coder-v2* For the vl model, it runs really well at a 23% CPU to 77% GPU ratio. Totally fine for my purposes. As for deepseek, it is a 40/60 ratio, but I luck out as it is a mixture of expert's model but even with the ratio, it is extremely performant. Gemma is by far my best model, and I have the most context room available at around 16k whereas the remainder I have sitting at 8k. Both gemma and qwen3.5 fit entirely in my GPUs VRAM. A couple things I noticed: Gemma4, is so good. Doesn't overthink, understands the prompt, remains as concise with the right tone I want. A really good day to day general model to work with. I also love the extra headroom for the context. Qwen3.5, a heavy thinker. Whilst it does a great job on the output, it spends a lot of time thinking and generating a lot of tokens. Power usage is pretty good, broke around 203W at one point and anything below that it just sat at whatever the power limit was set to. Qwen3-vl, also a major over-thinker. It spends so much time thinking that it balloons the context. I probably do not understand how to use this very well because when reading its thoughts it knows the answer pretty early on but it just gets into a thought trap (eh-yo). It does always output the correct answer, or the best it can, but I might move away from a reasoning vision model and stick to traditional ocr. If you have a better model or know how to prompt this better, I would love the help. Oh, one final note, this model LOVES power. Always maxes whatever I have, not that it increased performance directly, but it just loved power. Deepseek-coder-v2, this model rocks. It is extremely performant even though I technically on paper can't fit it. Especially for smaller asks with good bounds in place, it doesn't think, it just does and gives me excellent code back. I have yet to make it build me anything bigger but that is something I will experiment more with later. It is a weird one though, consistently using a fraction of the power budget available to it. Under 150W power limits, not once did my fans kick in on my GPU. Even my CPU fans (which are mucho loudo) rarely turned on, or if they did, they were not sounding like rocket engines. Not sure if those two are related but, eh, just something I noticed. Deepseek I think had some anomalous results with some spikes, but I can't be arsed to do them again. For the most part the results are fairly consistent and show a trend. Same for the vision model by qwen, oh well. My thoughts? It probably doesn't matter too much for most of us on a budget. Just let her rip, but if you want to shave off some heat, just lower your power down a little bit and monitor your temps. For the most part, you are probably fine. Honestly, it is 3am at this point, have a look at the spreadsheet! It was a lot of fun (I think) doing this. Interesting observations were made where I can balance my power limits, save... well pennies, and not have to listen to fans. So, works for me. https://docs.google.com/spreadsheets/d/1CEAr40nemlsK727QMvBnPcJDg548s-Bx/edit?usp=sharing&ouid=105501696463520933058&rtpof=true&sd=true Managers love graphs   submitted by   /u/bulletrhli [link]   [comments]