Llama.cpp and new model releases ...or why Great is the enemy of Good in the LLM world

Wait 5 sec.

llama.cpp is fundamental to this community. I remember when I went from struggling with transformers for a new model to just loading the model with llama.cpp with a change in how many layers ended up on the CPU. The phrase "Good is the enemy of great" is a central thesis from Jim Collins' 2001 book Good to Great... The idea being 'it is easy to settle for something that is merely adequate.' I would argue Llama.cpp holds fast to the slogan "Good is the enemy of Great", and not without good reason. When I have made this sort of complaint before, I was chastised about how my mindset and viewpoint would create technical debt challenges that could kill the project. So why continue arguing for my viewpoint? Llama.cpp in its effort to be sustainable is making unsustainable choices, at least for the masses. New software inference projects are gaining visibility and focus solely because they are not waiting for Great, but settling for Good enough... And the difference between Good enough and Great shouldn't stop the a release. A great example is GLM 5.3 Flash: On August 26th, release day, we had GLM 5.3 Flash with zero day support inside Unsloth Desktop built off Llama.cpp. EXL3 added support September 1st. Now, one month later, we still do not have support for GLM 5.3 Flash in Llama.cpp main. If 1 year is 7 years for dogs, what would 1 month be for LLMs? Major labs release models every 2 to 3 months on average. For some models, they will have little to no usage at all with llama.cpp because they are overshadowed by the next model release. Now some would say just use Unsloth then... Or EXL3. That supports my point. Llama.cpp is slowly dooming its widespread usage if everyone adopts that mentality. Others, more technically minded, would say just use a fork until it's fully released. This isn't just about me. There are too many using Ollama, LM Studio, KoboldCPP, or some other prebuilt binary to benefit from that suggestion. The point being made is convenience coupled with the pace of model releases will result in models not being used, or other platforms supplanting llama.cpp. Llama.cpp has 1.6k pull requests that sit waiting for the masses. Some or many likely don't deserve the light of day. But people have turned to solutions like DwarfStar just for specific model support. TLDR; I'm not arguing that Llama.cpp should throw caution to the wind and adopt every PR immediately, but it seems a different release process is needed. A user excited to use GLM 5.3 Flash shouldn’t have to learn how to fork and build software to continue using llama.cpp with the new model. Not to say the main branch should have this chaos, but a beta branch or one off binary releases could help. When the lead time from a functional version to the final release is over a month, it seems the energy to have a separate build with tentative GGUFs seems it is worth it. Unsloth clearly thinks so, and they're smarter than I. What do you think? If you agree, an upvote would be appreciated. Perhaps this will get the visibility needed to effect change with the creators of llama.cpp. If you don't, a comment explaining what I'm not considering, or a suggestion on how this could happen with less disruption would be valued...   submitted by   /u/silenceimpaired [link]   [comments]