COMMITS
July 22, 2026
M
[Attn] Support for heterogeneous hybrid-attention configurations (#1051)
Michael R committed
C
[CI] add ascend-a2-benchmark-ci and fix ascend-a2-ci (#1052)
ChunyuWei committed
July 21, 2026
C
[Ops] Add triton-ascend backend for KDA kernels (#1047)
ChunyuWei committed
Z
[Cleanup] Remove Ascend chunk_fwd_o kernels superseded by #1049 (#1050)
Zhiyuan Li committed
A
S
[Perf] Fuse ds in Ascend chunk_bwd_dqkwg to avoid recomputing do@v.T (#1048)
sunyi0505 committed
July 20, 2026
A
[Perf] Add opt-in TileLang RWKV6 intra kernel (#1045)
Annie Guo committed
S
July 19, 2026
M
[Perf] Cache and bucket Wall attention autotuning (#1041)
morluto committed
M
[Fix] Support value-split Wall backward with local deltas (#1042)
morluto committed
July 18, 2026
M
M
[Fix] Avoid concurrent LSE stores in value-split attention (#1033)
morluto committed
July 17, 2026
M
[Fix] Store split attention decode outputs at the correct offset (#1031)
morluto committed
S
[Perf] Optimize triton-ascend L2Norm with row tiling on Ascend NPU (#1036)
sunyi0505 committed
July 16, 2026
M
[Fix] fall back to Triton when TileLang has no usable nvcc compiler (#1026)
Matt Van Horn committed
H
July 15, 2026
S
July 12, 2026
C
[Ops] Add GDN kernels for triton-ascend backend (#1011)
ChunyuWei committed
July 9, 2026
Z
[Test] Skip flash-attn-dependent tests when flash-attn is not installed (#1014)
Zhiyuan Li committed
Z
Add get_max_length to FLA cache layer for transformers 5.x (#1009)
Zhiyuan Li committed
July 8, 2026
N
[Fix] Correct Mamba-3 decay parameterization (#1012)
Netanel Haber committed
July 7, 2026
Y
[Ops] Add Gluon backend for AttnRes (#1010)
Yu Zhang committed
Y
[Fix] Expose cp_context/disable_recompute in chunk_rwkv7 wrapper (#1004)
Yoake_0829 committed
Z
[Fix] avoid SymInt stride reads in activations under torch.compile (#1008)
Zhiyuan Li committed
X
[GDN] Add FlashQLA backend dispatch (#998)
Xi Lin committed
July 6, 2026
S
K
[Fix] drop deprecated tl.extra.cuda.libdevice prefix (#1005)
Kashif Rasul committed
July 3, 2026
C