COMMITS
July 22, 2026
S
[ET-VK][build] Fix Windows Vulkan/Android build and llama ETDump profiling
Stephen Jia committed
S
Revert "Revert "Qualcomm AI Engine Direct - improve llama3.2 3B TPS"" (#21150)
Scott Roy committed
S
Allow transposed convolution weights with a non-default dim order (#21035)
Suryansh Sijwali committed
J
migrate 3 fbcode-only TARGETS dirs with placeholder BUCK siblings (#21066)
Jon Janzen committed
J
Relax SmolLM2 QNN SQNR threshold (#21147)
Jacob Szwejbka committed
S
Revert "Qualcomm AI Engine Direct - improve llama3.2 3B TPS" (#21143)
Scott Roy committed
Q
Y
Record LLM model execution latency (#21040)
yenhao committed
S
Use dynamic CPU count for cmake --build -j in docs and test scripts (#20436)
Shamsudeen Saleem committed
R
Cortex-M backend: lower quantized GELU via the activation LUT (#21025)
RJ Ascani committed
B
Arm backend: Document RIFE accuracy handover (#21059)
Baris committed
R
Arm backend: Support SNORM RGBA runtime format (#20964)
Rob Elliott committed
A
Fuse consecutive clamps into a single clamp (#21013)
Andrew committed
S
Fix viable strict updating job (#21111)
Scott Roy committed
S
Flakey CI fix (#21110)
Scott Roy committed
July 23, 2026
K
Add functional update_and_attend KV-cache op (#21112)
Kiymet Akdemir committed
B
Cortex-M: Support avg_pool2d ceil_mode lowering (#21039)
Baris committed
S
Arm backend: Update transformers package to 5.3.0 (#21060)
SaoirseARM committed
P
step==1 contiguous fast path in portable compute_slice (#19606)
pssrawat committed
J
Reland two-pass channelwise gated delta rule kernel + OSS-safe bench BUCK (#21105)
Jacob Stevens committed
July 21, 2026
S
[ET-VK][sdpa] Vendor-adaptive head_dim output-tiling (TILE_N4=2) in GQA AV coop-GEMV
Stephen Jia committed
S
[ET-VK][sdpa] Reuse shared V cache across GQA query heads in AV coop-GEMV
Stephen Jia committed
S
[ET-VK][sdpa] Add SDPA operator perf benchmark binary (test_sdpa)
Stephen Jia committed
Q
Qualcomm AI Engine Direct - improve llama3.2 3B TPS (#20903)
qti-chenweng committed
P
[ExecuTorch][WebGPU] Fuse attention QKV projections in prefill (#21107)
pytorchbot committed
D
Preserve BatchNorm running state in composable quantizer
Deniz Kilinc committed
S
Revert "Arm backend: Rerun duplicate-user fusion after TOSA lowering" (#21106)
Scott Roy committed
S
Make extension/cuda:caller_stream fbcode-only
Scott Roy committed
J
gate fbcode-only TARGETS dirs with is_fbcode() and rename to BUCK
Jon Janzen committed
S
Add cooperative matrix dispatch for quantized linear
Sicheng Stephen Jia committed