COMMITS
September 10, 2026
K
P
Revert "[TRITONGPU] Preserve empty inner loops during flattening" (#11686)
peterbell10 committed
M
[TRITONGPU] Preserve empty inner loops during flattening (#11607)
Mingfei Guo committed
J
M
[Analysis] Preserve shared-memory aliases through calls (#11671)
Mario Lezcano Casado committed
M
[NFC] Remove dead synchronization-analysis paths (#11672)
Mario Lezcano Casado committed
J
[Gluon] Add per-thread inline assembly with descriptor operands (#11623)
jeffniu-openai committed
C
[ConSan] Slice proxy and barrier bookkeeping by proven indices (#11642)
Chenkai Mao committed
J
[triton_kernels] Add explicit nearest-even NVFP4 scale rounding (#11668)
Joey Yu committed
T
Update Blackwell ptxas to 13.4.59 (#11669)
Thomas Raoux committed
September 9, 2026
K
[CONSAN] Make consan and membar passes independent (#11597)
Keren Zhou committed
P
[ConSan] Stream tracking clears after barrier phase completion (#11452)
pawelszczerbuk committed
P
[ConSan] Stream large visibility and tracking clears (#11451)
pawelszczerbuk committed
P
[ConSan] Slice tracking updates for unique barrier slots (#11450)
pawelszczerbuk committed
P
[ConSan] Load read visibility for the issuing observer (#11449)
pawelszczerbuk committed
P
[ConSan] Slice proven single-buffer read visibility updates (#11448)
pawelszczerbuk committed
P
[ConSan] Use narrow masks for small logical thread sets (#11447)
pawelszczerbuk committed
P
[ConSan] Use masked stores for scratch overwrites (#11446)
pawelszczerbuk committed
K
[PROTON] Move Proton lowering before warp specialization (#11523)
Keren Zhou committed
L
[AMD] Fence LDS around wave synchronization (#11651)
Lei Zhang committed
P
[Backend] Ignore fractional cache policy for cp.async + respect cache modifier (#11660)
peterbell10 committed
S
[AMD] Fix additive strides for overlapping register bases (#11566)
Saeid Rostami committed
P
Revert "Infer slice encodings for reshapes" (#11658)
peterbell10 committed
P
M
[BACKEND] Allow warpshuffles when warp and block can read off broadcasted copies (#11646)
Mario Lezcano Casado committed
P
[Backend] Fix atomic_load fence placement (#11650)
peterbell10 committed
P
[Build] Update internal cuda download path (#11649)
peterbell10 committed
J
M
[BACKEND] Add native upcast fp4 to bf16/f16 (#11647)
Mario Lezcano Casado committed
P
[Language] Add dedicated atomic_{load,store} ops (#11591)
peterbell10 committed