-
Notifications
You must be signed in to change notification settings - Fork 4.1k
All issues
Issue creation is restricted in this repository
- #29831 · tianleiwu opened
on Jul 23, 2026
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#29853 In microsoft/onnxruntime;
[CPU] MatMulNBits accuracy_level=4 (int8 activation quant) selects wrong argmax token on massive-activation LLMs (Qwen3-0.6B, Phi-3.5-mini)
quantizationissues related to quantizationissues related to quantizationStatus: Open.#29849 In microsoft/onnxruntime;[Feature Request] [WebGPU EP] Enable more graph fusions
ep:WebGPUort-web webgpu providerort-web webgpu providerfeature requestrequest for unsupported feature or enhancementrequest for unsupported feature or enhancementplatform:webissues related to ONNX Runtime web; typically submitted using templateissues related to ONNX Runtime web; typically submitted using templateStatus: Open.#29841 In microsoft/onnxruntime;ORT 1.28.0 Release Candidates Available for Testing
.NETPull requests that update .net codePull requests that update .net codeStatus: Open.#29831 In microsoft/onnxruntime;CUDA AveragePool fails with CUDNN_STATUS_NOT_SUPPORTED when batch dimension reaches 65536
ep:CUDAissues related to the CUDA execution providerissues related to the CUDA execution providerStatus: Open.#29819 In microsoft/onnxruntime;[Pilot] Two-level (partition-time + kernel-instance) workspace estimation for MatMulNBits (CUDA EP, fpA_intB/CUTLASS)
ep:CUDAissues related to the CUDA execution providerissues related to the CUDA execution providerStatus: Open.#29810 In microsoft/onnxruntime;- Status: Open.#29809 In microsoft/onnxruntime;
[WebGPU EP] Deterministic audio corruption in Kokoro-82M TTS on Intel Iris Xe — fp32 corrupts a subset of inputs, fp16 corrupts all; same model clean on WASM EP and on Qualcomm Adreno
ep:WebGPUort-web webgpu providerort-web webgpu providermodel:transformerissues related to a transformer model: BERT, GPT2, Hugging Face, Longformer, T5, etc.issues related to a transformer model: BERT, GPT2, Hugging Face, Longformer, T5, etc.platform:mobileissues related to ONNX Runtime mobile; typically submitted using templateissues related to ONNX Runtime mobile; typically submitted using templateplatform:webissues related to ONNX Runtime web; typically submitted using templateissues related to ONNX Runtime web; typically submitted using templateStatus: Open.#29807 In microsoft/onnxruntime;- Status: Open.#29790 In microsoft/onnxruntime;
[Web] Race condition in Dawn build in Linux CI workflows
ep:WebGPUort-web webgpu providerort-web webgpu providerplatform:webissues related to ONNX Runtime web; typically submitted using templateissues related to ONNX Runtime web; typically submitted using templateStatus: Open.#29785 In microsoft/onnxruntime;[Build] Windows x64 QNN CI Pipeline and all Windows binskim builds fail with MSB8040 (missing Spectre-mitigated libs) on Win2022-GPU-A10 pool
buildbuild issues; typically submitted using templatebuild issues; typically submitted using templateep:QNNissues related to QNN exeution providerissues related to QNN exeution providerplatform:windowsissues related to the Windows platformissues related to the Windows platformStatus: Open.#29784 In microsoft/onnxruntime;Feature request: quantized (int8/fp8) KV cache in GroupQueryAttention with past/present shared buffer + CUDA graph capture
ep:CUDAissues related to the CUDA execution providerissues related to the CUDA execution providerquantizationissues related to quantizationissues related to quantizationStatus: Open.#29783 In microsoft/onnxruntime;