Making large AI models cheaper, faster and more accessible
COMMITS
/ examples/inference/benchmark_ops/benchmark_flash_decoding_attention.py May 14, 2024
S
add paged-attetionv2: support seq length split across thread block (#5707)
Steve Luo committed
May 5, 2024
Y
[Fix] Fix & Update Inference Tests (compatibility w/ main)
Yuanheng Zhao committed
April 30, 2024
April 25, 2024
S
[Inference/Kernel] Optimize paged attention: Refactor key cache layout (#5643)
Steve Luo committed
April 18, 2024