u/Hairy_Goose9089

Scaling my LLM inference for reply suggestions using disaggregated prefill
▲ 15 r/Vllm+1 crossposts

Scaling my LLM inference for reply suggestions using disaggregated prefill

Sharing my learning from separated prefill and decode into separate stages to increase the processing throughput and reducing TTFT significantly

saraswatmks.github.io
u/Hairy_Goose9089 — 6 days ago