Was getting bored trying to squeeze every last t/s out of my local model on my hardware, so I made a tps counter for my fingers instead (with real tokenizers, of course). My best is around 2 t/s. According to the page, that beats a 70B on a laptop CPU and is roughly 76x slower than an 8B on a 4090. So yeah. If you want something to do while waiting for your LLM to answer, give it a try! https://homoagens.github.io/human-tps/ PS: you can also share your result.. who knows, maybe I'll get famous :)   submitted by   /u/HomoAgens1 [link]   [comments]