Your current environment
Current Environment:
Collection
Collecting environment information...
==============================
OS : Ubuntu 20.04.6 LTS (x86_64)
GCC version : (Ubuntu 9.4.0-1ubuntu1~20.04.2) 9.4.0
Clang version : Could not collect
CMake version : version 4.1.0
Libc version : glibc-2.31
==============================
PyTorch Info
PyTorch version : 2.11.0+cu129
Is debug build : False
CUDA used to build PyTorch : 12.9
ROCM used to build PyTorch : N/A
XPU used to build PyTorch : N/A
==============================
Python Environment
Python version : 3.10.9 | packaged by conda-forge | (main, Feb 2 2023, 20:20:04) [GCC 11.3.0] (64-bit run
time)
Python platform : Linux-5.15.0-1084-aws-x86_64-with-glibc2.31
==============================
CUDA / GPU Info
Is CUDA available : True
CUDA runtime version : 12.1.105
CUDA_MODULE_LOADING set to :
GPU models and configuration : GPU 0: NVIDIA L4
Nvidia driver version : 570.133.07
cuDNN version : Could not collect
HIP runtime version : N/A
MIOpen runtime version : N/A
Is XNNPACK available : True
==============================
CPU Info
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Byte Order: Little Endian
Address sizes: 48 bits physical, 48 bits virtual
CPU(s): 8
On-line CPU(s) list: 0-7
Thread(s) per core: 2
Core(s) per socket: 4
Socket(s): 1
NUMA node(s): 1
Vendor ID: AuthenticAMD
CPU family: 25
Model: 1
Model name: AMD EPYC 7R13 Processor
Stepping: 1
CPU MHz: 2649.998
BogoMIPS: 5299.99
Hypervisor vendor: KVM
Virtualization type: full
L1d cache: 128 KiB
L1i cache: 128 KiB
L2 cache: 2 MiB
L3 cache: 16 MiB
NUMA node0 CPU(s): 0-7
Vulnerability Gather data sampling: Not affected
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Mmio stale data: Not affected
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed: Not affected
Vulnerability Spec rstack overflow: Mitigation; safe RET
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Retpolines; IBPB conditional; IBRS_FW; STIBP always-on; RSB filling; P
BRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Not affected
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mm
x fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good nopl nonstop_tsc cpuid extd_apicid
aperfmperf tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand
hypervisor lahf_lm cmp_legacy cr8_legacy abm sse4a misalignsse 3dnowprefetch topoext invpcid_single ssbd ibrs ibpb stib
p vmmcall fsgsbase bmi1 avx2 smep bmi2 invpcid rdseed adx smap clflushopt clwb sha_ni xsaveopt xsavec xgetbv1 clzero xsa
veerptr rdpru wbnoinvd arat npt nrip_save vaes vpclmulqdq rdpid
==============================
Versions of relevant libraries
[pip3] flashinfer-python==0.6.13
[pip3] numpy==2.2.6
[pip3] nvidia-cublas==13.1.0.3
[pip3] nvidia-cublas-cu12==12.9.1.4
[pip3] nvidia-cuda-cccl==13.3.3.4.1
[pip3] nvidia-cuda-cccl-cu12==12.9.27
[pip3] nvidia-cuda-crt==13.3.73
[pip3] nvidia-cuda-cupti==13.0.85
[pip3] nvidia-cuda-cupti-cu12==12.9.79
[pip3] nvidia-cuda-nvcc==13.2.86
[pip3] nvidia-cuda-nvcc-cu12==12.9.86
[pip3] nvidia-cuda-nvrtc==13.0.88
[pip3] nvidia-cuda-nvrtc-cu12==12.9.86
[pip3] nvidia-cuda-runtime==13.0.96
[pip3] nvidia-cuda-runtime-cu12==12.9.79
[pip3] nvidia-cuda-tileiras==13.2.86
[pip3] nvidia-cudnn-cu12==9.17.1.4
[pip3] nvidia-cudnn-cu13==9.19.0.56
[pip3] nvidia-cudnn-frontend==1.26.0
[pip3] nvidia-cufft==12.0.0.61
[pip3] nvidia-cufft-cu12==11.4.1.4
[pip3] nvidia-cufile==1.15.1.6
[pip3] nvidia-cufile-cu12==1.14.1.1
[pip3] nvidia-curand==10.4.0.35
[pip3] nvidia-curand-cu12==10.3.10.19
[pip3] nvidia-cusolver==12.0.4.66
[pip3] nvidia-cusolver-cu12==11.7.5.82
[pip3] nvidia-cusparse==12.6.3.3
[pip3] nvidia-cusparse-cu12==12.5.10.65
[pip3] nvidia-cusparselt-cu12==0.7.1
[pip3] nvidia-cusparselt-cu13==0.8.0
[pip3] nvidia-cutlass-dsl==4.5.2
[pip3] nvidia-cutlass-dsl-libs-base==4.5.2
[pip3] nvidia-cutlass-dsl-libs-cu13==4.5.2
[pip3] nvidia-ml-py==13.610.43
[pip3] nvidia-nccl-cu12==2.28.9
[pip3] nvidia-nccl-cu13==2.28.9
[pip3] nvidia-nvjitlink==13.0.88
[pip3] nvidia-nvjitlink-cu12==12.9.86
[pip3] nvidia-nvshmem-cu12==3.4.5
[pip3] nvidia-nvshmem-cu13==3.4.5
[pip3] nvidia-nvtx==13.0.85
[pip3] nvidia-nvtx-cu12==12.9.79
[pip3] nvidia-nvvm==13.2.86
[pip3] pyzmq==27.1.0
[pip3] tokenspeed-triton==3.8.10.post20260709
[pip3] torch==2.11.0+cu129
[pip3] torch_c_dlpack_ext==0.1.5
[pip3] torchaudio==2.11.0+cu129
[pip3] torchcodec==0.15.0+cu129
[pip3] torchvision==0.26.0+cu129
[pip3] transformers==5.14.1
[pip3] triton==3.6.0
[conda] numpy 2.2.6 pypi_0 pypi
[conda] nvidia-cublas-cu12 12.1.3.1 pypi_0 pypi
[conda] nvidia-cuda-cupti-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-cuda-nvrtc-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-cuda-runtime-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-cudnn-cu12 8.9.2.26 pypi_0 pypi
[conda] nvidia-cufft-cu12 11.0.2.54 pypi_0 pypi
[conda] nvidia-curand-cu12 10.3.2.106 pypi_0 pypi
[conda] nvidia-cusolver-cu12 11.4.5.107 pypi_0 pypi
[conda] nvidia-cusparse-cu12 12.1.0.106 pypi_0 pypi
[conda] nvidia-nccl-cu12 2.20.5 pypi_0 pypi
[conda] nvidia-nvjitlink-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-nvtx-cu12 12.1.105 pypi_0 pypi
[conda] torch 2.3.0+cu121 pypi_0 pypi
[conda] torchvision 0.18.0 pypi_0 pypi
[conda] transformers 4.57.0 pypi_0 pypi
[conda] triton 2.3.0 pypi_0 pypi
==============================
vLLM Info
ROCM Version : Could not collect
vLLM Version : 0.25.1
vLLM Build Flags:
CUDA Archs: Not Set; ROCm: Disabled; XPU: Disabled
GPU Topology:
GPU0 CPU Affinity NUMA Affinity GPU NUMA ID
GPU0 X 0-7 0 N/A
Legend:
X = Self
SYS = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)
NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node
PHB = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)
PXB = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)
PIX = Connection traversing at most a single PCIe bridge
NV# = Connection traversing a bonded set of # NVLinks
==============================
Environment Variables
LD_LIBRARY_PATH=/home/ubuntu/diffusion/.venv/lib/python3.10/site-packages/cv2/../../lib64:/opt/amazon/efa/lib:/opt/amazo
n/openmpi/lib:/opt/aws-ofi-nccl/lib:/usr/local/cuda/lib:/usr/local/cuda:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUP
TI/lib64:/usr/local/cuda/targets/x86_64-linux/lib:/usr/local/lib:/usr/lib
PYTORCH_NVML_BASED_CUDA_CHECK=1
TORCHINDUCTOR_COMPILE_THREADS=1
TORCHINDUCTOR_CACHE_DIR=/tmp/torchinductor_ubuntu
VLLM_WORKER_MULTIPROC_METHOD=spawn
Summary of my setup:
- vLLM 0.25.1 (
vllm-0.25.1+cu129-cp38-abi3-manylinux_2_28_x86_64.whl from the v0.25.1 GitHub release)
- torch 2.11.0+cu129
- Triton 3.6.0 (torch's pinned version) — also reproduced on Triton 3.7.1
- GPU: NVIDIA L4 (24GB, SM 8.9 / Ada)
- Driver 570.133.07 (CUDA 12.8)
- Python 3.10, Ubuntu 20.04
- Model:
RedHatAI/diffusiongemma-26B-A4B-it-NVFP4 (architecture DiffusionGemmaForBlockDiffusion)
🐛 Describe the bug
🐛 Describe the bug
Serving a DiffusionGemma model fails during engine startup inside kernel_warmup()
(vllm/model_executor/warmup/kernel_warmup.py, line 47), which unconditionally imports
minimax_m3_msa_warmup regardless of which architecture is loaded.
That import chain reaches vllm/models/minimax_m3/common/ops/index_topk.py, where Triton's
@triton.jit decorator on _topk_index_merge_kernel fails while parsing the kernel source:
File ".../triton/runtime/jit.py", line 470, in __init__
src = src[re.search(r"^def\s+\w+\s*\(", src, re.MULTILINE).start():]
AttributeError: 'NoneType' object has no attribute 'start'
There appear to be two separate problems:
- Triton cannot parse the
index_topk.py kernel source in this environment. Reproduced on
both Triton 3.6.0 (the version torch 2.11.0 pins) and Triton 3.7.1.
- The import is unconditional, so a broken kernel belonging to one architecture prevents
serving any model. kernel_warmup() starts with this import before any architecture check,
meaning a MiniMax-specific problem becomes a total serving outage for unrelated models. This
import should be conditional on the loaded architecture, or wrapped defensively.
Everything before this point succeeds: weights load, torch.compile completes, CUDA graphs capture,
and the KV cache is allocated. The failure is the last step before the API server would bind.
Reproduction
VLLM_USE_V2_MODEL_RUNNER=1 vllm serve RedHatAI/diffusiongemma-26B-A4B-it-NVFP4 \
--trust-remote-code \
--max-num-seqs 4 \
--max-model-len 32768 \
--hf-overrides '{"diffusion_sampler": "entropy_bound", "diffusion_entropy_bound": 0.1}' \
--default-chat-template-kwargs '{"enable_thinking": true}'
Workaround
Wrapping the import with a no-op stub allows the server to start and serve normally:
try:
from vllm.model_executor.warmup.minimax_m3_msa_warmup import (
minimax_m3_msa_warmup,
)
except Exception:
def minimax_m3_msa_warmup(*args, **kwargs):
pass
With this patch applied, the same command reaches Application startup complete. and serves
requests correctly.
Possibly related
The MiniMax M3 roadmap (#45668) and the recent replacement of these Triton kernels with
tok_sparse_select from MSA suggest this code path is being actively reworked, so the parse
failure may already be moot on main. The unconditional import in kernel_warmup() still affects
released versions and would be worth a defensive fix regardless.
Full traceback
EngineCore startup traceback (click to expand)
Traceback (most recent call last):
File ".../vllm/v1/engine/core.py", line 1200, in run_engine_core
engine_core = EngineCoreProc(*args, engine_index=dp_rank, **kwargs)
File ".../vllm/tracing/otel.py", line 178, in sync_wrapper
return func(*args, **kwargs)
File ".../vllm/v1/engine/core.py", line 966, in __init__
super().__init__(
File ".../vllm/v1/engine/core.py", line 133, in __init__
kv_cache_config = self._initialize_kv_caches(vllm_config)
File ".../vllm/tracing/otel.py", line 178, in sync_wrapper
return func(*args, **kwargs)
File ".../vllm/v1/engine/core.py", line 321, in _initialize_kv_caches
self.model_executor.initialize_from_config(kv_cache_configs)
File ".../vllm/v1/executor/abstract.py", line 124, in initialize_from_config
compilation_times: list[CompilationTimes] = self.collective_rpc(
File ".../vllm/v1/executor/uniproc_executor.py", line 92, in collective_rpc
result = run_method(self.driver_worker, method, args, kwargs)
File ".../vllm/v1/serial_utils.py", line 510, in run_method
return func(*args, **kwargs)
File ".../vllm/tracing/otel.py", line 178, in sync_wrapper
return func(*args, **kwargs)
File ".../vllm/v1/worker/gpu_worker.py", line 758, in compile_or_warm_up_model
kernel_warmup(self)
File ".../vllm/model_executor/warmup/kernel_warmup.py", line 47, in kernel_warmup
from vllm.model_executor.warmup.minimax_m3_msa_warmup import (
File ".../vllm/model_executor/warmup/minimax_m3_msa_warmup.py", line 7, in <module>
from vllm.models.minimax_m3.nvidia.model import MiniMaxM3SparseAttention
File ".../vllm/models/minimax_m3/__init__.py", line 17, in <module>
from .nvidia.model import (
File ".../vllm/models/minimax_m3/nvidia/model.py", line 72, in <module>
from vllm.models.minimax_m3.common.indexer import MiniMaxM3Indexer
File ".../vllm/models/minimax_m3/common/indexer.py", line 39, in <module>
from vllm.models.minimax_m3.common.ops.index_topk import (
File ".../vllm/models/minimax_m3/common/ops/__init__.py", line 5, in <module>
from .index_topk import (
File ".../vllm/models/minimax_m3/common/ops/index_topk.py", line 549, in <module>
def _topk_index_merge_kernel(
File ".../triton/runtime/jit.py", line 923, in decorator
return JITFunction(
File ".../triton/runtime/jit.py", line 756, in __init__
super().__init__(fn)
File ".../triton/runtime/jit.py", line 469, in __init__
src = src[re.search(r"^def\s+\w+\s*\(", src, re.MULTILINE).start():]
AttributeError: 'NoneType' object has no attribute 'start'
Before submitting a new issue...
Your current environment
Current Environment:
Collection
Collecting environment information...
==============================
OS : Ubuntu 20.04.6 LTS (x86_64)
GCC version : (Ubuntu 9.4.0-1ubuntu1~20.04.2) 9.4.0
Clang version : Could not collect
CMake version : version 4.1.0
Libc version : glibc-2.31
==============================
PyTorch Info
PyTorch version : 2.11.0+cu129
Is debug build : False
CUDA used to build PyTorch : 12.9
ROCM used to build PyTorch : N/A
XPU used to build PyTorch : N/A
==============================
Python Environment
Python version : 3.10.9 | packaged by conda-forge | (main, Feb 2 2023, 20:20:04) [GCC 11.3.0] (64-bit run
time)
Python platform : Linux-5.15.0-1084-aws-x86_64-with-glibc2.31
==============================
CUDA / GPU Info
Is CUDA available : True
CUDA runtime version : 12.1.105
CUDA_MODULE_LOADING set to :
GPU models and configuration : GPU 0: NVIDIA L4
Nvidia driver version : 570.133.07
cuDNN version : Could not collect
HIP runtime version : N/A
MIOpen runtime version : N/A
Is XNNPACK available : True
==============================
CPU Info
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Byte Order: Little Endian
Address sizes: 48 bits physical, 48 bits virtual
CPU(s): 8
On-line CPU(s) list: 0-7
Thread(s) per core: 2
Core(s) per socket: 4
Socket(s): 1
NUMA node(s): 1
Vendor ID: AuthenticAMD
CPU family: 25
Model: 1
Model name: AMD EPYC 7R13 Processor
Stepping: 1
CPU MHz: 2649.998
BogoMIPS: 5299.99
Hypervisor vendor: KVM
Virtualization type: full
L1d cache: 128 KiB
L1i cache: 128 KiB
L2 cache: 2 MiB
L3 cache: 16 MiB
NUMA node0 CPU(s): 0-7
Vulnerability Gather data sampling: Not affected
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Mmio stale data: Not affected
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed: Not affected
Vulnerability Spec rstack overflow: Mitigation; safe RET
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Retpolines; IBPB conditional; IBRS_FW; STIBP always-on; RSB filling; P
BRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Not affected
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mm
x fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good nopl nonstop_tsc cpuid extd_apicid
aperfmperf tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand
hypervisor lahf_lm cmp_legacy cr8_legacy abm sse4a misalignsse 3dnowprefetch topoext invpcid_single ssbd ibrs ibpb stib
p vmmcall fsgsbase bmi1 avx2 smep bmi2 invpcid rdseed adx smap clflushopt clwb sha_ni xsaveopt xsavec xgetbv1 clzero xsa
veerptr rdpru wbnoinvd arat npt nrip_save vaes vpclmulqdq rdpid
==============================
Versions of relevant libraries
[pip3] flashinfer-python==0.6.13
[pip3] numpy==2.2.6
[pip3] nvidia-cublas==13.1.0.3
[pip3] nvidia-cublas-cu12==12.9.1.4
[pip3] nvidia-cuda-cccl==13.3.3.4.1
[pip3] nvidia-cuda-cccl-cu12==12.9.27
[pip3] nvidia-cuda-crt==13.3.73
[pip3] nvidia-cuda-cupti==13.0.85
[pip3] nvidia-cuda-cupti-cu12==12.9.79
[pip3] nvidia-cuda-nvcc==13.2.86
[pip3] nvidia-cuda-nvcc-cu12==12.9.86
[pip3] nvidia-cuda-nvrtc==13.0.88
[pip3] nvidia-cuda-nvrtc-cu12==12.9.86
[pip3] nvidia-cuda-runtime==13.0.96
[pip3] nvidia-cuda-runtime-cu12==12.9.79
[pip3] nvidia-cuda-tileiras==13.2.86
[pip3] nvidia-cudnn-cu12==9.17.1.4
[pip3] nvidia-cudnn-cu13==9.19.0.56
[pip3] nvidia-cudnn-frontend==1.26.0
[pip3] nvidia-cufft==12.0.0.61
[pip3] nvidia-cufft-cu12==11.4.1.4
[pip3] nvidia-cufile==1.15.1.6
[pip3] nvidia-cufile-cu12==1.14.1.1
[pip3] nvidia-curand==10.4.0.35
[pip3] nvidia-curand-cu12==10.3.10.19
[pip3] nvidia-cusolver==12.0.4.66
[pip3] nvidia-cusolver-cu12==11.7.5.82
[pip3] nvidia-cusparse==12.6.3.3
[pip3] nvidia-cusparse-cu12==12.5.10.65
[pip3] nvidia-cusparselt-cu12==0.7.1
[pip3] nvidia-cusparselt-cu13==0.8.0
[pip3] nvidia-cutlass-dsl==4.5.2
[pip3] nvidia-cutlass-dsl-libs-base==4.5.2
[pip3] nvidia-cutlass-dsl-libs-cu13==4.5.2
[pip3] nvidia-ml-py==13.610.43
[pip3] nvidia-nccl-cu12==2.28.9
[pip3] nvidia-nccl-cu13==2.28.9
[pip3] nvidia-nvjitlink==13.0.88
[pip3] nvidia-nvjitlink-cu12==12.9.86
[pip3] nvidia-nvshmem-cu12==3.4.5
[pip3] nvidia-nvshmem-cu13==3.4.5
[pip3] nvidia-nvtx==13.0.85
[pip3] nvidia-nvtx-cu12==12.9.79
[pip3] nvidia-nvvm==13.2.86
[pip3] pyzmq==27.1.0
[pip3] tokenspeed-triton==3.8.10.post20260709
[pip3] torch==2.11.0+cu129
[pip3] torch_c_dlpack_ext==0.1.5
[pip3] torchaudio==2.11.0+cu129
[pip3] torchcodec==0.15.0+cu129
[pip3] torchvision==0.26.0+cu129
[pip3] transformers==5.14.1
[pip3] triton==3.6.0
[conda] numpy 2.2.6 pypi_0 pypi
[conda] nvidia-cublas-cu12 12.1.3.1 pypi_0 pypi
[conda] nvidia-cuda-cupti-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-cuda-nvrtc-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-cuda-runtime-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-cudnn-cu12 8.9.2.26 pypi_0 pypi
[conda] nvidia-cufft-cu12 11.0.2.54 pypi_0 pypi
[conda] nvidia-curand-cu12 10.3.2.106 pypi_0 pypi
[conda] nvidia-cusolver-cu12 11.4.5.107 pypi_0 pypi
[conda] nvidia-cusparse-cu12 12.1.0.106 pypi_0 pypi
[conda] nvidia-nccl-cu12 2.20.5 pypi_0 pypi
[conda] nvidia-nvjitlink-cu12 12.1.105 pypi_0 pypi
[conda] nvidia-nvtx-cu12 12.1.105 pypi_0 pypi
[conda] torch 2.3.0+cu121 pypi_0 pypi
[conda] torchvision 0.18.0 pypi_0 pypi
[conda] transformers 4.57.0 pypi_0 pypi
[conda] triton 2.3.0 pypi_0 pypi
==============================
vLLM Info
ROCM Version : Could not collect
vLLM Version : 0.25.1
vLLM Build Flags:
CUDA Archs: Not Set; ROCm: Disabled; XPU: Disabled
GPU Topology:
GPU0 CPU Affinity NUMA Affinity GPU NUMA ID
GPU0 X 0-7 0 N/A
Legend:
X = Self
SYS = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)
NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node
PHB = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)
PXB = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)
PIX = Connection traversing at most a single PCIe bridge
NV# = Connection traversing a bonded set of # NVLinks
==============================
Environment Variables
LD_LIBRARY_PATH=/home/ubuntu/diffusion/.venv/lib/python3.10/site-packages/cv2/../../lib64:/opt/amazon/efa/lib:/opt/amazo
n/openmpi/lib:/opt/aws-ofi-nccl/lib:/usr/local/cuda/lib:/usr/local/cuda:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUP
TI/lib64:/usr/local/cuda/targets/x86_64-linux/lib:/usr/local/lib:/usr/lib
PYTORCH_NVML_BASED_CUDA_CHECK=1
TORCHINDUCTOR_COMPILE_THREADS=1
TORCHINDUCTOR_CACHE_DIR=/tmp/torchinductor_ubuntu
VLLM_WORKER_MULTIPROC_METHOD=spawn
Summary of my setup:
vllm-0.25.1+cu129-cp38-abi3-manylinux_2_28_x86_64.whlfrom the v0.25.1 GitHub release)RedHatAI/diffusiongemma-26B-A4B-it-NVFP4(architectureDiffusionGemmaForBlockDiffusion)🐛 Describe the bug
🐛 Describe the bug
Serving a DiffusionGemma model fails during engine startup inside
kernel_warmup()(
vllm/model_executor/warmup/kernel_warmup.py, line 47), which unconditionally importsminimax_m3_msa_warmupregardless of which architecture is loaded.That import chain reaches
vllm/models/minimax_m3/common/ops/index_topk.py, where Triton's@triton.jitdecorator on_topk_index_merge_kernelfails while parsing the kernel source:There appear to be two separate problems:
index_topk.pykernel source in this environment. Reproduced onboth Triton 3.6.0 (the version torch 2.11.0 pins) and Triton 3.7.1.
serving any model.
kernel_warmup()starts with this import before any architecture check,meaning a MiniMax-specific problem becomes a total serving outage for unrelated models. This
import should be conditional on the loaded architecture, or wrapped defensively.
Everything before this point succeeds: weights load, torch.compile completes, CUDA graphs capture,
and the KV cache is allocated. The failure is the last step before the API server would bind.
Reproduction
VLLM_USE_V2_MODEL_RUNNER=1 vllm serve RedHatAI/diffusiongemma-26B-A4B-it-NVFP4 \ --trust-remote-code \ --max-num-seqs 4 \ --max-model-len 32768 \ --hf-overrides '{"diffusion_sampler": "entropy_bound", "diffusion_entropy_bound": 0.1}' \ --default-chat-template-kwargs '{"enable_thinking": true}'Workaround
Wrapping the import with a no-op stub allows the server to start and serve normally:
With this patch applied, the same command reaches
Application startup complete.and servesrequests correctly.
Possibly related
The MiniMax M3 roadmap (#45668) and the recent replacement of these Triton kernels with
tok_sparse_selectfrom MSA suggest this code path is being actively reworked, so the parsefailure may already be moot on
main. The unconditional import inkernel_warmup()still affectsreleased versions and would be worth a defensive fix regardless.
Full traceback
EngineCore startup traceback (click to expand)
Before submitting a new issue...