SIGN IN SIGN UP

Qualcomm AI Engine Direct - Add block quantization to llama (#10225)

- Add CLI argument to use block quantization for llama

Co-authored-by: Chun-I Tsai <chunit@qti.qualcomm.com>
C
Chun-I Tsai committed
e6c7b30c39f94d6dd3a74fa5a9e8d50ac6eb3334
Parent: 8b4500b
Committed by GitHub <noreply@github.com> on 4/16/2025, 7:03:54 PM