vLLM vs SGLang — Benchmarking LLM Serving Engines on an RTX 4090
I benchmarked vLLM and SGLang on a single RTX 4090, serving meta-llama/Meta-Llama-3-8B-Instruct in fp16. I wanted to see how the two leading open-source serving engines compare under identical load...