COMMITS
July 29, 2026
V
Stop quantize_ from leaking global fp32 matmul precision (complete fix) (#4643)
Vasiliy Kuznetsov committed
V
Remove disabled A100 perf workflow and dead llama-benchmark scripts (#4641)
Vasiliy Kuznetsov committed
V
Split PT2E tests into dedicated CPU and GPU CI workflows (#4644)
Vasiliy Kuznetsov committed
V
Pin CI to PyTorch 2.13 for release; drop unused torchaudio (#4638)
Vasiliy Kuznetsov committed
V
Fix ROCm rowwise-fp8 warning crash (#4640)
Vasiliy Kuznetsov committed
July 27, 2026
V
add warnings for float8 rowwise inference on CUDA 12.9 (#4588)
Vasiliy Kuznetsov committed
July 26, 2026
L
[pat] Fix SVD class leakage into `nn.Linear` from #4587 (#4613)
Lisa Jin committed
July 24, 2026
July 23, 2026
L
[pat] Low-rank from internal repo (#4586)
Lisa Jin committed
Z
Flydsl mxfp8 quantize (#4357)
Zachary Streeter committed
A
Add workflow to triage main CI failures with Claude (#4467)
andrewor14 committed
July 22, 2026
T
J
Fix Int4WeightFakeQuantizer and Float8FakeQuantizer missing enabled attribute (#4336)
Javier De Jesus committed
S
Fixes `cutlass.base_dsl._mlir_helpers` import failure with nvidia-cutlass-dsl 4.6.1 (#4590)
Syed Tousif Ahmed committed
July 20, 2026
N
kernel: synchronize autotuner cache updates (#4554)
Neil Schemenauer committed
N
kernel: avoid partial Triton lazy initialization (#4553)
Neil Schemenauer committed
B
docs: fix Sphinx toctree doc build warning for performant_kernels.rst (#3863) (#4515)
Brittney Lilly committed
July 17, 2026
T
Revert "Revert "pt2e: preserve mutable buffer inputs during prepare""
Tom Allsop committed
July 14, 2026
S
Gate int8 channelwise neondot qmatmul behind runtime FEAT_DotProd check
Scott Roy committed
July 10, 2026
N
optim: cache qmap values as tuples (#4552)
Neil Schemenauer committed
N
cpu: lock ukernel registration tables (#4561)
Neil Schemenauer committed
N
Re-enable building under a free-threaded Python interpreter (#4537)
Neil Schemenauer committed
July 9, 2026
A
[xpu][float8] device agnostic test base,compile, numerics_integration (#3823)
Artur Lesniak committed
July 8, 2026
P
Fix typo in e2e model level benchmarks - h100->b200 (#4547)
Pqlet committed
July 7, 2026
O
fix: use `Int4WeightOnlyConfig.group_size` in fake quant configs (#4518)
Oriol Liñan committed
V
skip failing CI tests (#4549)
Vasiliy Kuznetsov committed
July 6, 2026
N
float8: snapshot FSDP precomputed scale (#4555)
Neil Schemenauer committed
J
[ROCm] Add HIP device synchronize event to profiler overhead filter (#4450)
Jagadish Krishnamoorthy committed