COMMITS
August 15, 2026
D
Y
[GFX950] Relocate MLA Gluon kernel and unify decode dispatch (#4450)
yinfengLiu committed
M
S
[dtype] Map FP8 to torch.float8_e4m3fn on RDNA3 (#4764)
skysnow2001 committed
N
[GFX1250] [Gluon] Add multicast support for MoE a8w4 (#4562)
Nicholas Susanto committed
August 14, 2026
L
[triton][moe] Add sigmoid score_mode to the routing top-k (#4688)
lijinpei-amd committed
A
fused_qk_rope_reshape_and_cache kernel optimizations (gfx950) (#4719)
amirumoAMD committed
F
D
gfx1250 opus gemm splitk fuse (#4246)
demonsan committed
A
CI: auto-update split test FILE_TIMES (#4649)
aiter-gh-app[bot] committed
A
[module_mla_metadata*] refactor and remove torch (#4729)
amd-ruitang3 committed
A
[aiterTensor::empty] remove use in .cu (#4724)
amd-ruitang3 committed
M
docs: fix ISA kernel optimization guide and example scripts (#4561)
mario-ant committed
X
ci: pin PyTorch test image (#4744)
Xin Huang committed
Y
Y
[MLA] Support 48-head 128-dim reduction (#4727)
yinfengLiu committed
J
feat: add param for combine quant (#4746)
JiaoliangYu committed
X
CI: add always-run Aiter test gate (#4418)
Xin Huang committed
F
Add Opus hd192 hybrid buffer path for large KV (>4GiB). (#4473)
fangche123 committed
Y
[opus_moe] Unify A8W4 Stage2 with a runtime-K decode pipeline (#4723)
yifehuan committed
August 13, 2026
J
[triton] fix sparse decode gathering zeros from a strided cache (#4673)
jiacao-amd committed
S
[Triton] add a copilot review instructions and modify readme file (#4720)
Satya Nikhil Kodukula committed
S
[CI] Skip Triton test suites on docs-only changes (#4734)
Satya Nikhil Kodukula committed
J
[Triton] Optimize chunk_delta_attn performance. (#4683)
jianhao committed
J
Avoid duplicate fp32 output alloc in torch_moe_stage2 reference (#4717)
Johannes Graner committed
J
CI: publish per-case vLLM DI accuracy and enable Kimi (#4722)
JiaoliangYu committed
X
Restore default GitHub Pages docs deployment (#4728)
Xin Huang committed
J
feat: support fp32 chunk states in GDN prefill (#4366)
junna2016 committed
B
[MI355] add 8wave pipeline to a8w8 bpreshuffle gemm (#4714)
BingYuan.Zhou committed
A
[fix] add missing <optional> header in topk_plain_kernels.cu (#4725)
amd-ruitang3 committed