Netflix Details Its In-House LLM Serving Platform with Triton and vLLM

Wait 5 sec.

Netflix has described the production lessons behind bringing LLM inference into its internal serving platform, including the challenges of supporting different model sizes, hardware requirements, and rapidly evolving inference engines. By Matt Foster