COMMITS
November 20, 2025
J
update readme
Jiayi Yuan committed
J
Fix README formatting and update release notes
Jiayi Yuan committed
J
Merge pull request #47 from QuiverDance/fix-flashattn-prefill-mask
Jiayi Yuan committed
Q
fix: flash attention prefill mask for padded batches
QuiverDance committed
September 25, 2025
Z
Update README.md
Zirui Liu committed
Z
Update README.md
Zirui Liu committed
Z
Update README.md
Zirui Liu committed
January 19, 2025
October 10, 2024
J
Merge pull request #29 from condy0919/triton
Jiayi Yuan committed
J
Merge pull request #30 from jy-yuan/dependabot/pip/gradio-5.0.0
Jiayi Yuan committed
D
Bump gradio from 3.35.2 to 5.0.0
dependabot[bot] committed
September 26, 2024
C
Make KIVI quantization work on cuda:{1,2,...}
condy committed
August 27, 2024
J
add template for llama3
jy-yuan committed
August 23, 2024
Z
Update gemv_cuda.cu
Zirui Liu committed
Z
Update gemv_cuda.cu
Zirui Liu committed
August 16, 2024
Z
Update README.md
Zirui Liu committed
June 16, 2024
Z
Merge pull request #21 from yifeikong/main
Zirui Liu committed
June 14, 2024
Y
Add missing flash_attn_func import in llama_kivi model
Yifei Kong committed
June 7, 2024
Z
Update README.md
Zirui Liu committed
May 24, 2024
Z
Update README.md
Zirui Liu committed
April 28, 2024
J
add memory and speed test code
jy-yuan committed
April 17, 2024
Z
Update pyproject.toml
Zirui Liu committed
Z
Fix the bug of prompt template
Zirui Liu committed
April 12, 2024
J
update readme
jy-yuan committed
J
upd long bench readme
jy-yuan committed
J
add results for longchat and mistral
jy-yuan committed
J
add support for mistral and longchat
jy-yuan committed