-
Notifications
You must be signed in to change notification settings - Fork 831
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Treat total_num_pages as a physical KV cache budget, not the page-table size
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3557
opened Sep 21, 2026 by
0z5a
Loading…
Add an opt-in batch-invariant BF16 GEMM
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3556
opened Sep 21, 2026 by
0z5a
Loading…
Add an opt-in float32 reduction dtype for the row-parallel all-reduce
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3555
opened Sep 21, 2026 by
0z5a
Loading…
Reuse the host split sizes GroupedLinear's forward already computed
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3554
opened Sep 21, 2026 by
0z5a
Loading…
Isolate FlashAttention 2/3 imports so a broken install cannot break TE
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3553
opened Sep 21, 2026 by
0z5a
Loading…
Accumulate delayed wgrads into shared (tied) parameters
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3552
opened Sep 21, 2026 by
0z5a
Loading…
Support learnable softmax (attention sink) with cp_comm_type="p2p"
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3551
opened Sep 20, 2026 by
0z5a
Loading…
Fix FusedSGD zero_grad argument forwarding
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3549
opened Sep 19, 2026 by
wujingyue
Contributor
Loading…
4 tasks done
[PyTorch] Enable LayerNormLinear UB AG overlap under no_grad
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3546
opened Sep 18, 2026 by
ravimajeti
Loading…
7 of 8 tasks
[PyTorch] Keep DistributedWeight objects on ctx across saved-tensor hooks
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3545
opened Sep 18, 2026 by
xrennvidia
Collaborator
Loading…
1 of 13 tasks
feat: vmm slot for local cuda graph offload
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3544
opened Sep 18, 2026 by
Xdydy
Loading…
1 of 13 tasks
Use PyTorch CUTLASS for BF16 grouped GEMM on SM100
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3542
opened Sep 18, 2026 by
wujingyue
Contributor
Loading…
5 of 7 tasks
[PyTorch][CP] Support compact metadata with FA4
#3540
opened Sep 17, 2026 by
sudhakarsingh27
Member
Loading…
13 tasks
fix: add missing NVTE_BHSD labels to fused_attn to_string
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3538
opened Sep 17, 2026 by
andrewwhitecdw
Contributor
Loading…
Prototype of green context + VMM localization
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
test(attention): anchor the context-parallel suite to an independent reference
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3536
opened Sep 17, 2026 by
nvegesna-netizen
Contributor
Loading…
[PyTorch] [torch.compile] torch.compile support for LayerNormLinear and LayerNormMLP
#3534
opened Sep 17, 2026 by
pggPL
Collaborator
Loading…
[PyTorch] Retain NVFP4 RNG tensors through quantization
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3533
opened Sep 17, 2026 by
Connor-XY
Loading…
fix(attention): stop handing FlashAttention 4 the -1 window sentinel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3532
opened Sep 17, 2026 by
nvegesna-netizen
Contributor
Loading…
[PyTorch] Extend no-load-balance CP to the a2a comm type
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3530
opened Sep 16, 2026 by
Rudin6
Loading…
13 tasks
feat(attention): cuDNN FROST attention backend for head_dim in (256, 512], with context parallelism
2.21.0
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3527
opened Sep 16, 2026 by
nvegesna-netizen
Contributor
Loading…
[Pytorch][NCCL EP] Create tokens_per_expert on pinned CPU memory when running eager
#3526
opened Sep 16, 2026 by
YangFei1990
Collaborator
Loading…
8 of 13 tasks
Previous Next
ProTip!
Updated in the last three days: updated:>2026-09-18.