SIGN IN SIGN UP

Add gguf q4_k quantization (#2001)

* Add gguf q4_k_s quantization

Summary:
Didn't implement the algorithm to choose_qparams from gguf, since it's complicated, e.g. https://github.com/ggml-org/llama.cpp/blob/f423981ac806bf031d83784bcb47d2721bc70f97/ggml/src/ggml-quants.c#L744 and https://github.com/ggml-org/llama.cpp/blob/f423981ac806bf031d83784bcb47d2721bc70f97/ggml/src/ggml-quants.c#L827C14-L827C28

but implemented a simple choose_qparams that can fit the gguf format:
Q4_K: w = q * block_scale(6-bit) + block_min(6-bit)

Test Plan:
python test/prototype/test_gguf_quant.py

Reviewers:

Subscribers:

Tasks:

Tags:

* fix

* test with phi4

* pre-commit run

* update

* run precommit

* format
J
Jerry Zhang committed
ef10f348d739df8d239173c830fa9bc5aaa816ae
Parent: 5802d2d
Committed by GitHub <noreply@github.com> on 4/8/2025, 6:06:32 PM