Yup you read the headline right, “Yet another inference engine” although I have a twist for you! this isn’t faster than Llama.cpp :D It started this spring when I wanted to build my own agentic solution, I have limited hardware (A M1 Mac with 8 GB RAM + A Gaming Machine I7-7700 16GB RAM +3070 8 GB VRAM that I use mostly with Moonlight to game on the mac) so I needed something that could work with smaller models and also making sure it could “digest” whatever I threw at it. Naturally I hit multiple walls since smaller models are dumb as shit and often enough end up losing their context window and or just not answering at all. Having configured my agentic solution to also include a JSON parser and trying to get it to work with both llama.cpp and Ollama I finally grew tired of building the dependencies outside of the inference engine, and this summer Xyntetik-Runner was born (yeah Xyntetik… here we go again). C was the language of choice because why not, I had extremely low experience in Writing code overall, C seemed like the right choice mostly because ever since I’ve been growing up, anything competent needs to be in C, not sure if it’s true but it’s been hammered over and over again with me so it’s kind of stuck… The main issue here I was trying to solve was the truncated tool calling issue I had but also, I didn’t need it to support multiple models, so I made it work good enough with the models I was working with the most. Then something opened up a bit, I got access to one of my friends AI Machines (A Nvidia Blackwell Card where I got a 24 GB MIG slice) and suddenly I could actually start using some more competent local models and I think this is where the spark began… I wanted to see if we could do more on smaller hardware, not necessarily faster (and not 0,00003 tokens/s either) but what are the “challenges” if you will. So, what did I actually Focus on: - Truncated tool calling, A response comes back, no mess no fuss no features needed in between, it returns a valid Json - Schema enforcement, no more invalid tokens so no parse and retry loop, this is a killer for agentic workflows btw… - Runner doesn’t eat memory when idle, I can have the engine “on” on my laptop and only when it calls the model the RAM gets eaten - Runner Can Train LoRa directly on a 4 Bit file I use, no FP16 copy needed. - Runner is Token Identical to Llama.cpp, tested it on Gemma4-Moe and got token for token certification - Since I have three different OS/HW there’s not just metal, Cuda, intel + amd CPU aswell, whetever that gives. - OpenAI compatible server, since I needed it to be a --serve endpoint its included. Little did I realize how deep this rabbit hole would be so well... I ended up adding features as I needed them, took a great interest in the challenges of modern AI (I’ve learned so much since then, and yet it still feels sometimes I know nothing). This is where I realized I Needed Lora Adapters, Checksum verifications, Processes/kill switches for testing, cadence, theses. You name it, I probably got some embryo somewhere among my Terabytes of testing grounds. And at the same time, I wanted the agentic solution I’m developing to grow so whatever features needed for that got added Aswell. The direction then is two-fold: enterprise support for verified inference locally and the other track, Research. (Some of you might remember my post on Xyntetik-Kvist-14B, that is exactly what came out of this, and yes, I learned my lesson there, Fable got nothing to do with this post, this is all me so it’s your own fault for getting less facts and more rambling! :D) The Suite/enterprise part is still under construction, and I hope to have something there that can actually be of use to the industry. Runner will however be free forever (*cough* Apache 2.0 *Cough*) since I think the world needs this kind of things, the world might not need Runner specifically but it’s important that we all try to drive this evolution forward. Research is an interesting topic, on my HF I am currently posting more and more on the current branches around “machine without human” Called Genesis, exploring How machine2Machine language works, what happens if there’s no human teacher or language in the loop? Interesting read if you have the time and it’s an active branch where time is the only factor on when results get published. The best part, I use Runner for all my work, so I dogfood a lot, which means bugs, features and so forth gets patched and fixed as soon as they appear. Would love it to get some input, feedback, forks or whatever, happy to help, happy to evolve, or just shut up if you want me to… And the links: Runner: https://github.com/Joakimpalm-Zen/xyntetik-runner HF: https://huggingface.co/Joakimpalm-Zen Main Web: https://xyntetik.com/ Runner is developed with the Assistance of, Astra, Fable, Opus, Sol and all the other fine “people” that we usually deal with. And last but not least, tired as hell now, going to sleep, let me know if there’s anything, or nothing, or something….   submitted by   /u/ZenZombie117 [link]   [comments]