I've been working on Modly, an open-source desktop app that turns images or prompt into 3D meshes with only local models. It has a chat agent that can operate the app. In v0.4.3 I made llama.cpp the default engine and built the agent around it. Why llama.cpp - I wanted direct control over how the model runs: context size, GPU offload, KV-cache quantization, flash attention. - Plain GGUF files. Pick from a small catalog, or drop any .gguf into the models folder and it shows up. How it runs - One llama-server process per loaded model, on localhost only. - You can keep several models loaded at once. The default count is sized from your VRAM, and idle servers get unloaded so 3D generation has room. - The agent is a standard OpenAI-style tool-calling loop against the app's own API: read mesh info, decimate, smooth, list/run/create workflows, unload models from VRAM, etc. - The model library shows size, quant and an estimated VRAM footprint, and grades each model on tool calling. Grades are marked as either measured with a small eval suite in the app or estimated from public benchmarks, so you know which is which. What the video shows Qwen 3.5 4B Q4_K_M on an RTX 3060 12 GB. I ask it to cut a 2.6M-triangle mesh down to 300k. It calls `decimate_mesh` with the right path and target and reports the result. About 9 s with the model already loaded; the first call takes ~40 s because llama-server has to start and load the weights. Honest limitations - Small models sometimes misreport results. In one test the decimation stopped above the target (UV seams limit how far it can simplify), and the model made up a reason instead of just reporting the number. I'm thinking about feeding the tool output back more explicitly. - It's an assistant on top of the app, not a replacement for the UI. Multi-step workflow creation is noticeably less reliable at 4B than single tool calls. - Other backends are still optional: any OpenAI-compatible endpoint works, including your own llama-server. Local llama.cpp is the default, and nothing leaves your machine unless you configure something else. Question for you Which models up to 8B have you found most reliable for tool calling on llama.cpp? Qwen 3 4B / 3.5 4B work best for me so far. GPT-OSS 20B is good but too heavy next to a 3D generation model on 12 GB. Also curious whether people would rather tune the llama-server flags themselves or keep sane defaults.   submitted by   /u/Lightnig125 [link]   [comments]