-
-
Notifications
You must be signed in to change notification settings - Fork 19.9k
All issues
Issue creation is restricted in this repository
- #42770 · WoosukKwon opened
on May 15, 2026 21 - #44280 · BugenZhao opened
on Jun 2, 2026 20 - #48168 · simon-mo opened
on Jul 9, 2026 3
Issues
is:issue state:open
is:issue state:open
Search results
[Perf] #48137 costs ~10.6% spec-decode acceptance and #48660 shifts output distributions on DeepSeek-V4-Flash — isolated via #48660-only arm on a production 2-node deployment
cpuRelated to CPU backendsRelated to CPU backendsStatus: Open.#49927 In vllm-project/vllm;[Bug]: EngineDeadError NVFP4 marlin
bugSomething isn't workingSomething isn't workingStatus: Open.#49926 In vllm-project/vllm;[Bug][XPU]: GDN attention silently corrupts memory under load — fix merged in vllm-xpu-kernels but requirements/xpu.txt pins a release that predates it
bugSomething isn't workingSomething isn't workingintel-gpuRelated to Intel GPURelated to Intel GPUStatus: Open.#49924 In vllm-project/vllm;[Bug]: [Regression] Assertion res == CUresult::CUDA_SUCCESS failed in FlashMLA (phase1.cuh) for DeepSeek-V4 on v0.26.0 (Works in v0.25.0)
bugSomething isn't workingSomething isn't workingStatus: Open.#49922 In vllm-project/vllm;- Status: Open.#49921 In vllm-project/vllm;
[Bug]: DiffusionGemma - Unconditional minimax_m3 warmup import in kernel_warmup() crashes engine startup for unrelated models (Triton JIT fails to parse index_topk kernel)
bugSomething isn't workingSomething isn't workingStatus: Open.#49920 In vllm-project/vllm;- Status: Open.#49905 In vllm-project/vllm;
[Bug]: Tiered KV offload promotes every waiting request (which fills primary DRAM pool)
bugSomething isn't workingSomething isn't workingStatus: Open.#49902 In vllm-project/vllm;- Status: Open.#49896 In vllm-project/vllm;
[Bug]: SpeculativeConfig method="draft_model" cannot load mixed-precision compressed-tensors checkpoints (config_groups)
bugSomething isn't workingSomething isn't workingStatus: Open.#49893 In vllm-project/vllm;[Bug]: GLM-5.2-NVFP4 produces garbled/incorrect output and hits NotImplementedError in forward_mha on GB10 (SM121a) with FLASHINFER_MLA_SPARSE_SM120
bugSomething isn't workingSomething isn't workingStatus: Open.#49886 In vllm-project/vllm;