-
Notifications
You must be signed in to change notification settings - Fork 2.6k
Pull requests: NVIDIA/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Feat/golden prairie multimodal
#17043
opened Jul 30, 2026 by
WeiHaocheng
Collaborator
Loading…
1 task
[NVBUG-6448152][test] TEST ONLY: legacy GEN selection and blocking discriminator
#17042
opened Jul 30, 2026 by
chienchunhung
Collaborator
•
Draft
[https://nvbugs/6525010][fix] Fix TestLlama3_3_70BInstruct::test_nvfp…
#17040
opened Jul 30, 2026 by
liji-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6462303][chore] Unwaive KV cache v2 scheduler tests with executor loop tracing
#17039
opened Jul 30, 2026 by
yizhang-nv
Member
Loading…
5 tasks done
[#17022][fix] trtllm-eval: make --max_output_length reachable on aime25/aime26
#17037
opened Jul 30, 2026 by
thorjohnsen
Collaborator
•
Draft
1 task done
[#17021][fix] DeepSeek-V4: merge request tools into the existing system message
#17036
opened Jul 30, 2026 by
thorjohnsen
Collaborator
•
Draft
1 task done
[#17020][fix] DeepSeek-V4: honor thinking_budget in apply_chat_template
#17035
opened Jul 30, 2026 by
thorjohnsen
Collaborator
•
Draft
1 task done
[Bug]: TensorRT-LLM 1.3.0rc18 Llama-3-70B inference failure on 8x NVIDIA RTX 6000D
#17034
opened Jul 30, 2026 by
pjdurden
Loading…
[TRTLLM-14779][fix] Clear capture-only sampling override from cached CUDA graph metadata
#17033
opened Jul 29, 2026 by
xwang233
Collaborator
Loading…
[NVBUG-6327718][test] Unwaive test_disaggregated_videomme[nemotron_na…
#17031
opened Jul 29, 2026 by
aswinvisva
Collaborator
•
Draft
1 task
[TRTLLM-14778][perf] Enable CUDA graphs for encoder-decoder encoder steps
#17030
opened Jul 29, 2026 by
pranav-nvidia
Contributor
•
Draft
1 task
[None][feat] Delegate MX loading to ModelExpress strategies
#17029
opened Jul 29, 2026 by
zhengluo-nv
Loading…
1 task
[#17027][fix] Realign misaligned int32 index slices in FlashInfer GDN decode
#17028
opened Jul 29, 2026 by
Navjot10
Loading…
1 task done
[None][fix] Size attention context workspace with the real cross-KV length
#17026
opened Jul 29, 2026 by
pranav-nvidia
Contributor
•
Draft
1 task done
[None][fix] Size the trtllm-gen multi-CTA KV counter buffer for beam search
#17014
opened Jul 29, 2026 by
pranav-nvidia
Contributor
•
Draft
3 of 4 tasks
[TRTLLM-14622][feat] Add generic VisualGen static LoRA config
VisualGen
#17012
opened Jul 29, 2026 by
yibinl-nvidia
Collaborator
•
Draft
1 task
[None][fix] Use HND mapping for MiniMax-M3 MSA KV cache
#17011
opened Jul 29, 2026 by
peihu-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6525011][fix] Store the owning output tensor in _graph_output_refs for each graph's lifetime…
#17010
opened Jul 29, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[TRTLLM-14609][chore] Remove ENABLE_CONFIGURABLE_MOE escape hatch and remaining MoE legacy relics
#17009
opened Jul 29, 2026 by
xxi-nv
Collaborator
Loading…
4 tasks done
[None][feat] Add opt-in JSONL device-time logging to torch AutoTuner
#17007
opened Jul 29, 2026 by
chenfeiz0326
Collaborator
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.