COMMITS
July 30, 2026
C
[Refactor] Remove intrinsic compatibility facade (#2812)
Chaofan Lin committed
K
[BugFix] Resolve partial scalar reduce barrier participation (#2777)
KellyFrog committed
R
C
[CUDA] Extend the GEMM FMA fallback to SM75 (#2811)
Chennes committed
July 29, 2026
C
[Refactor] Remove unused tilelang.common package (#2810)
Chaofan Lin committed
C
[BugFix] Skip descriptor TMA for device-bound copy bases (#2803)
Chaofan Lin committed
C
T
[Fix] Refine architecture guards (#2790)
Tong WU committed
M
M
G
[Metal] Add line-level threadgroup qualifier scanning (pass 5) (#2796)
Gengyuan Bai committed
C
Y
[CUDA] Support arbitrary TMEM layouts (#2785)
Yongqi Zhuo committed
July 28, 2026
C
[FFI] Support apache-tvm-ffi 0.1.12 (#2795)
Chaofan Lin committed
L
L
[TIR][Language] Add typed vector lane extraction API (#2789)
Lei Wang committed
Y
[Metal] M5 Cooperative Tensor T.gemm (#2252)
Yichen Yan committed
G
[Metal] Add 16-byte alignment padding to shared/threadgroup memory (#2786)
Gengyuan Bai committed
B
[BugFix] Honor nan_propagate in reduce max/min/absmax clear=False write-back (#2788)
BoHao Chen committed
P
L
[CUDA][Transform] Fix PCWS index dtype handling (#2783)
Lei Wang committed
P
[BugFix] Preserve loop step when unrolling loops (#2784)
penguin_wwy committed
July 27, 2026
W
[BugFix] Fallback non-16B cluster bulk copies (#2683)
wcx committed
T
C
[TileOP] Add SM70 GEMM FMA fallback (#2339)
cklxx committed
J
C
[BugFix] Fix ROCm intrinsic resolution after language dialect refactor (#2779)
Chaofan Lin committed