u/thepetek

▲ 7 r/Vllm+1 crossposts

Qwen with cache offload vLLM

Has anyone gotten KV cache offloading working with Qwen on vLLM? No matter what configuration I try, I get errors and it crashes. I saw an old issue that Qwen arch is supported for offload in vLLM but that doesn’t seem right. Anyone have working settings they care to share?

reddit.com
u/thepetek — 1 day ago