Inference on one node is easy, but scaling to millions of users taught me about new things on a job

Wait 5 sec.

I know how exciting it is to play out different configurations of an LLM locally and improve speed, quality etc by custom optimisations. I was doing this for a long time, then I got to work on a job where I worked with inference providers scaling to millions of users. Then one node with multiple GPUs storing the weights vs several nodes sharing different users, caching, autoscaling, routing, all of these aren't something I could have spent money myself renting GPUs to learn. I am very thankful to have that job experience, and would love to share everything I learned about system design around inference. Also this is an extremely demanded skill and agents aren't quite there at automating it fully, but although you can still be a good mecha pilot. Here's a blog I wrote summarising the knowledge I gathered. I have written other blogs too, this is one of many: https://medium.com/@abhijithneilabraham/learning-llm-inference-from-a-single-gpu-to-millions-of-users-e4d5e216662d Feel free to reach out if any questions!   submitted by   /u/metalvendetta [link]   [comments]