-
Notifications
You must be signed in to change notification settings - Fork 2.8k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 2 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[Performance]: sm89 ada blockwise FP8 GEMM copies block-scale factors to shared memory 4x/128x redundantly (stride-0 scale TV layouts)
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentStatus: Open.#19566 In NVIDIA/TensorRT-LLM;[Bug] _load_bin_or_path_file hides the real torch.load error behind UnboundLocalError
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#19564 In NVIDIA/TensorRT-LLM;[Bug]: DisaggClusterManager drops remaining watch events when one event raises
bugSomething isn't workingSomething isn't workingDisaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Status: Open.#19551 In NVIDIA/TensorRT-LLM;[Bug]: Beam search with a small KV cache pool fails CUDA graph warmup with "No free block found"
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceStatus: Open.#19527 In NVIDIA/TensorRT-LLM;[Bug]: Whisper (PyTorch backend) fails on any repeated request when KV block reuse is on: "Request requires multimodal_embed_mask_cumsum for chunked prefill or KV-cache reuse"
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19515 In NVIDIA/TensorRT-LLM;[Bug]: cross Kv sizing with whisper
bugSomething isn't workingSomething isn't workingKV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19514 In NVIDIA/TensorRT-LLM;[Feature]: Add native text reranking support to the PyTorch backend
feature requestNew feature or request. This includes new model, dtype, functionality supportNew feature or request. This includes new model, dtype, functionality supportPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19510 In NVIDIA/TensorRT-LLM;- Status: Open.#19501 In NVIDIA/TensorRT-LLM;
[Bug]: LTX-2 asymmetric scale factors swap height and width
bugSomething isn't workingSomething isn't workingStatus: Open.#19495 In NVIDIA/TensorRT-LLM;[Bug]: AutoDeploy logger treats verbose and internal_error as info
AutoDeploy<NV> AutoDeploy Backend<NV> AutoDeploy BackendbugSomething isn't workingSomething isn't workingStatus: Open.#19491 In NVIDIA/TensorRT-LLM;[Bug]: Speculative decoding + TorchSampler: unseeded requests share one RNG stream → identical/degenerate samples at low concurrency
bugSomething isn't workingSomething isn't workingSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#19487 In NVIDIA/TensorRT-LLM;[New Model]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
new modelRequest to add a new modelRequest to add a new modelStatus: Open.#19481 In NVIDIA/TensorRT-LLM;