COMMITS
July 9, 2026
L
[CI] Fix cargo-deny config flag ordering (#48170)
Lucas Wilkinson committed
M
[ROCm] Revert Part of `[ROCm] Fix pooling startup workspace lock` #47912 (#48154)
Micah Williamson committed
L
[CI] Increase extract hidden states TP2 timeout (#48161)
Lucas Wilkinson committed
Z
[ROCm] Synchronize sparse MLA metadata before graph replay (#47404)
ZihaoMu committed
Z
[KV Connector][Mooncake] Apply SWA lookup mask before hashing/key build (#47317)
Zhewen Li committed
T
[ROCM][DSV32][Perf][MTP] Enable UNIFORM_BATCH CG mode in rocm_aiter_mla_sparse (#45149)
Teemu Virolainen committed
N
[Bugfix][MRV2] Reset num_accepted_tokens on add_request in all modes (#48132)
Nick Hill committed
W
L
[Bugfix] Preserve tensor causal metadata for grouped attention (#48135)
Lucas Wilkinson committed
C
[ROCm][CI] Set all timeout_in_minutes to 180 (#48146)
Charlie Fu committed
T
[Bugfix] Guard CUDA-only rms_norm_per_block_quant in FUSED_OPS for non-CUDA builds (#47296)
Tsvika Shapira committed
B
Pin PyNvVideoCodec to tested 2.0.4 wheel (#48056)
Brandon Pelfrey committed
K
[CI] Annotate built Docker image tags on the Buildkite build page (#48101)
Kevin H. Luu committed
W
H
Migrate Olmo and Olmo2 to the Transformers modeling backend (#48100)
Harry Mellor committed
T
Remove PersimmonForCausalLM and FuyuForCausalLM model architectures (#48096)
Tiezhen WANG committed
C
M
Sanitize server file paths from validation error responses (#46415)
Muhammad Fawaz committed
B
[Bugfix] Fix race condition in KVBlockZeroer (#48085)
Benjamin Chislett committed
W
Add Intel XPU Docker release pipeline (#47880)
wenjun liu committed
T
Remove TeleChatForCausalLM (#47989)
Tiezhen WANG committed
J
[Bugfix] Fix Qwen3-ASR transcription streaming postprocessing (#42478)
JooHo Lee committed
L
[CPU] Fix Qwen-Next SSM type for AMX GDN (#48073)
Li, Jiang committed
C
[KV Offloading] Add free block iterator for CPU offload scheduling (#47849)
Chauncey committed
C
[XPU][LoRA] Fix torch.compile DEVICE_LOST by avoiding view-mutation in LoRA shrink (#47944)
Chaojun Zhang committed
July 8, 2026
S
[ROCm][CI][MoE] Fix double-transpose of fused w3 expert weights (#47874)
stefankoncarevic committed
Q
N
[Bugfix] Use int8 workspace for FlashInfer MLA decode (#48046)
Nick Hill committed
H
Fix embed scaling + CUDA graphs in Transformers modelling backend (#48010)
Harry Mellor committed