Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Treat total_num_pages as a physical KV cache budget, not the page-table size community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3557 opened Sep 21, 2026 by 0z5a Loading…
Add an opt-in batch-invariant BF16 GEMM community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3556 opened Sep 21, 2026 by 0z5a Loading…
Add an opt-in float32 reduction dtype for the row-parallel all-reduce community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3555 opened Sep 21, 2026 by 0z5a Loading…
Reuse the host split sizes GroupedLinear's forward already computed community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3554 opened Sep 21, 2026 by 0z5a Loading…
Isolate FlashAttention 2/3 imports so a broken install cannot break TE community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3553 opened Sep 21, 2026 by 0z5a Loading…
Accumulate delayed wgrads into shared (tied) parameters community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3552 opened Sep 21, 2026 by 0z5a Loading…
Support learnable softmax (attention sink) with cp_comm_type="p2p" community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3551 opened Sep 20, 2026 by 0z5a Loading…
Fix FusedSGD zero_grad argument forwarding community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3549 opened Sep 19, 2026 by wujingyue Contributor Loading…
4 tasks done
KDA attention 2.20
#3548 opened Sep 18, 2026 by ksivaman Member Loading…
8 of 13 tasks
[PyTorch] Compile Linear and Bias operations with forward fusion
#3547 opened Sep 18, 2026 by pggPL Collaborator Draft
6 of 9 tasks
[PyTorch] Enable LayerNormLinear UB AG overlap under no_grad community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3546 opened Sep 18, 2026 by ravimajeti Loading…
7 of 8 tasks
[PyTorch] Keep DistributedWeight objects on ctx across saved-tensor hooks community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3545 opened Sep 18, 2026 by xrennvidia Collaborator Loading…
1 of 13 tasks
feat: vmm slot for local cuda graph offload community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3544 opened Sep 18, 2026 by Xdydy Loading…
1 of 13 tasks
Use PyTorch CUTLASS for BF16 grouped GEMM on SM100 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3542 opened Sep 18, 2026 by wujingyue Contributor Loading…
5 of 7 tasks
[PyTorch][CP] Support compact metadata with FA4
#3540 opened Sep 17, 2026 by sudhakarsingh27 Member Loading…
13 tasks
fix: add missing NVTE_BHSD labels to fused_attn to_string community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3538 opened Sep 17, 2026 by andrewwhitecdw Contributor Loading…
Prototype of green context + VMM localization community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3537 opened Sep 17, 2026 by WanZzzzzz Contributor Draft
13 tasks
test(attention): anchor the context-parallel suite to an independent reference community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3536 opened Sep 17, 2026 by nvegesna-netizen Contributor Loading…
[Docs] Refresh README
#3535 opened Sep 17, 2026 by sbhavani Collaborator Loading…
5 of 13 tasks
[PyTorch] Retain NVFP4 RNG tensors through quantization community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3533 opened Sep 17, 2026 by Connor-XY Loading…
fix(attention): stop handing FlashAttention 4 the -1 window sentinel community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3532 opened Sep 17, 2026 by nvegesna-netizen Contributor Loading…
[PyTorch] Extend no-load-balance CP to the a2a comm type community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3530 opened Sep 16, 2026 by Rudin6 Loading…
13 tasks
feat(attention): cuDNN FROST attention backend for head_dim in (256, 512], with context parallelism 2.21.0 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3527 opened Sep 16, 2026 by nvegesna-netizen Contributor Loading…
[Pytorch][NCCL EP] Create tokens_per_expert on pinned CPU memory when running eager
#3526 opened Sep 16, 2026 by YangFei1990 Collaborator Loading…
8 of 13 tasks
ProTip! Updated in the last three days: updated:>2026-09-18.