-
2026-07-12
-
2026-07-01
-
2026-06-09
vLLM solves head-of-line blocking by design. Most hardware can’t run vLLM. That gap has consequences — and they don’t show up in benchmarks on 128GB machines.
Writing on distributed systems, infrastructure, and ML serving.
vLLM solves head-of-line blocking by design. Most hardware can’t run vLLM. That gap has consequences — and they don’t show up in benchmarks on 128GB machines.