Inference Engines will become a series of one-offs

Wait 5 sec.

ninfer, dwarfstar, Splash, llamAmpere, gufo, etc. We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch. The list will continue to grow They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months. But new ones will take their place... THESIS One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower A few axioms you probably accept: the more general a codebase, the harder it is to cleanly fit new features in over time. This reduces the pace of innovation. The smaller, the faster AI coding is getting better and cheaper. Thus, the barrier to creating an inference engine is dropping many coding tasks are difficult to completely give to AI (or a human) because they are not fully specified. But "Make tok/s go up" in an inference engine fork for one hardware/model combo is fully specified, and is therefore a great candidate for 100% autonomous implementations to be perfectly fine in terms of quality (as long as correctness tests are included, which is trivial). No human bottleneck. String these axioms together, and I arrive at general engines, like llama.cpp, will perpetually lag behind these one-offs in development speed, and therefore token speed none of the one-offs will be able to maintain generality and dev speed over time one-off inference engines for specific hardware/model combinations will continue to proliferate, and be loved Thank you for coming to my ted talk IMPLICATIONS This thesis brings up an interesting question: What elements of inference engines WILL remain in common? Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago. Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc? One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it. There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt. Maybe we'll see more 'half-general' inference engines that just target one hardware platform. So still general on the dimension of models, but not on hardware. Splash could be an example. Nobody wants to continuously scan github/reddit/x for the best inference engine for their model/rig. Some will just have their agent custom make one. But I think a larger number will not do that. So, hardware-specific communities will form. Think r/appleM2Max32gbLLM and r/4090And64gbRamLLM, etc (however that actually ends up organizing. exaggerating a bit on the names.) ALTERNATIVE FUTURES Scenarios where the one-off future doesn't happen: General inference engines find a way to 'plugin-ify' the model/hardware specific kernels and compute graph so you can swap them at runtime. So you'd download not just a .gguf, but also an .inference_recipe to go with it, which contains the optimizations for your specific hardware, for that specific model. Maybe those optimizations will make it into llama.cpp mainline in 3 months, but you can use them today, without a fork. General inference engines find a way to AI-ify their workflow so much that they maintain quality and codebase coherence but also achieve the same development velocity for each model/hardware platform as the one-offs. I think this is the best for everyone involved. The full vision of something like MLIR, or Mojo is realized to a sufficient degree. ie writing hardware-optimized kernels is fully and invisibly done by compilers, no hardware-specific tinkering needed anymore for each silicon platform) Then, inference engines that cover all hardware/models would be much more manageable to maintain and add features to. btw, if you really want to have an impact, solve this. The world will thank you for centuries to come. Unfortunately not many people have even conceptualized this as a goal. P.S. there's growth in a dimension separate from single model/hardware engines which is more like "frontrunning a general inference engine's features because it's slower to pull in PRs". Freetoken, BeeLlama, etc. Not as model- or hardware- specific as the other examples I've given. Haven't thought much about that dimension.   submitted by   /u/netherreddit [link]   [comments]