COMMITS
July 28, 2026
M
MPK skills: agent skill suites for MPK development (#742)
Muheng Li committed
Y
Support Qwen3-32B: fix split-K linear grid for non-dividing K dims (#744)
Yizheng Jiao committed
S
adding inkling, GLM, optimized kimi dflash (#745)
Sina Lin committed
July 12, 2026
Y
Fix SM100 attention smem overflow blocking GQA >= 8:1 (#702) (#739)
Yizheng Jiao committed
June 24, 2026
Z
test: add MoE residual double-count regression test (null-residual path) (#731)
Zhanming (Jerry) Liang committed
June 22, 2026
M
Add Ferret frozen-gate kernel-agent system (3 subagents + skill) (#730)
Muheng Li committed
June 15, 2026
L
Support DFlash for Kimi-K2.6 (#728)
Letian Ruan committed
Z
pass format (#724)
Zhihao Jia committed
Z
Fix double-counted MoE residual under tensor parallelism (#722)
Zhanming (Jerry) Liang committed
Z
Fix moe_w13_linear_layer grid_dim for tensor parallelism (#723)
Zhanming (Jerry) Liang committed
June 13, 2026
C
perf: flatten the MLA kv-cache gather and add 4-way ILP (~2.3x on the gather) (#720)
Chris Fregly committed
C
fix: per-token FP8 quantize must not derive the batch row from blockIdx.x under MPK (#719)
Chris Fregly committed
June 10, 2026
Z
Update README.md for publication
Zhihao Jia committed
Z
fix: remove NVSHMEM_NO_DEVICE_LIB (#713)
Zephyr Zhao committed
Z
test: improve test mode interface (split from dpskv3) (#712)
Zephyr Zhao committed
June 9, 2026
Z
Feat/tensor view core (#711)
Zephyr Zhao committed
June 8, 2026
L
[MPK] Support EAGLE3 Speculative Decoding for Qwen3 (#692)
Letian Ruan committed
June 5, 2026
S
Add MPK batch size perf CI for Qwen3; pass MAX_TOKENS to SM100 attention (#709)
Shao Wang committed
June 4, 2026
M
Ferret dispatcher subagent (#701)
Muheng Li committed
S
fix: clamp TMA box dims for splitk linear on SM100 to fix deadlock (#705)
Shao Wang committed
June 3, 2026
S
ci: filter idle GPUs by memory usage and add job-level timeout (#706)
Shao Wang committed
June 2, 2026
S
Update Qwen3 CI workflow: conda env, idle GPU selection, cleanup (#703)
Shao Wang committed
May 31, 2026
S
fix: two regressions from PR #660 in offline persistent kernel (#704)
Shao Wang committed
May 27, 2026
S
[Online Serving] Continuous batching, streaming, and OpenAI server support (#660)
Shuaiwei Huang committed
S
fix: apply clang-format to all cpp files (#697)
Shao Wang committed
April 29, 2026
Z
fix: recover test mode (#673)
Zephyr Zhao committed
April 27, 2026
A
Add load_mpk for kernel reuse (#618)
Arav Tewari committed
April 26, 2026
Z
Parallel-path DAG compilation support in register_mugraph (#665)
Zephyr Zhao committed
M
quick bug fix of wrong qwen3 demo (#670)
Muheng Li committed
April 24, 2026
M
DeepSeek V3 on B200: FP8 + MLA TP + MLA prefill + MTP (#667)
Muheng Li committed