COMMITS
/ torchao/quantization/README.md March 5, 2025
A
Revert "Move torchao/_models to benchmarks/_models" (#1844)
Apurva Jain committed
March 4, 2025
A
Move torchao/_models to benchmarks/_models (#1784)
Apurva Jain committed
February 24, 2025
J
Update README.md (#1758)
Jerry Zhang committed
February 14, 2025
V
make quantize_.set_inductor_config None by default (#1716)
Vasiliy Kuznetsov committed
V
update torchao READMEs with new configuration APIs (#1711)
Vasiliy Kuznetsov committed
January 17, 2025
A
Update supported dtypes for fp8 (#1573)
Apurva Jain committed
January 9, 2025
J
Skip calling unwrap_tensor_subclass for torch 2.7+ (#1531)
Jerry Zhang committed
December 17, 2024
D
[WIP] Codebook quantization flow (#1299)
DerekLiu35 committed
M
Expose zero_point_domain as arguments (#1401)
Meng, Hengyu committed
December 16, 2024
H
gemlite integration in torchao (#1034)
HDCharles committed
November 29, 2024
S
Benchmark intel xpu (#1259)
Swift.Sun committed
November 27, 2024
S
Benchamarking (#1353)
Scott Roy committed
November 20, 2024
2
fix typo of README.md (#1318)
22dimensions committed
I
Fix pickle.dump missing file argument typo in README (#1316)
Ismayil Ismayilov committed
November 14, 2024
H
support W4A8 Marlin kernel (#1113)
HandH1998 committed
October 25, 2024
A
Move float8 out of prototype in quantization README (#1166)
Apurva Jain committed
October 15, 2024
A
Add generic fake quantized linear for QAT (#1020)
andrewor14 committed
October 10, 2024
A
Move and rename GranularityType -> Granularity (#1038)
andrewor14 committed
October 7, 2024
A
Dynamic Float8 benchmarking llama (#1017)
Apurva Jain committed
September 28, 2024
L
Fix WOQ int8 failures (#884)
leslie-fang-intel committed
September 18, 2024
H
Update README.md (#903)
HDCharles committed
September 17, 2024
V
move float8 inference README contents to prototype section (#901)
Vasiliy Kuznetsov committed
September 16, 2024
V
Update README.md for float8 inference (#896)
Vasiliy Kuznetsov committed
September 11, 2024
H
README and benchmark improvements (#867)
HDCharles committed
September 8, 2024
M
Update README (#823)
Mark Saroufim committed
September 6, 2024
J
Expose hqq through `uintx_weight_only` API (#786)
Jerry Zhang committed
D
Add sparse marlin AQT layout (#621)
Diogo Venâncio committed
September 5, 2024
J
Add uintx quant to generate and eval (#811)
Jerry Zhang committed
H
Update README.md (#814)
HDCharles committed
H
adding kv cache quantization to READMEs (#813)
HDCharles committed
H
int4 fixes and improvements (#804)
HDCharles committed
September 4, 2024
J
Add more information to quantized linear module and added some logs (#782)
Jerry Zhang committed
August 28, 2024
A
Update method names to support intx and floatx changes (#775)
Apurva Jain committed
August 25, 2024
M
1 more doc revamp (#745)
Mark Saroufim committed
August 16, 2024
J
Update README.md (#649)
Jerry Zhang committed
August 14, 2024
M
retry version guard fix (#679)
Mark Saroufim committed
August 7, 2024
R
Update README.md (#618)
Raziel committed
August 6, 2024
J
Update README.md (#619)
Jerry Zhang committed
August 3, 2024
J
Unpin nightly version (#593)
Jerry Zhang committed
August 1, 2024
J
Allow `benchmark_model` to accept args and kwargs (#586)
Jerry Zhang committed
July 26, 2024
J
Refactor LinearActQuantizedTensor (#542)
Jerry Zhang committed
July 10, 2024
J
Update calls to `quantize_` everywhere (#496)
Jerry Zhang committed
July 4, 2024
J
Renaming `quantize` to `quantize_` (#467)
Jerry Zhang committed
July 3, 2024
J
Add more docs for int4_weight_only API that targets tinygemm (#469)
Jerry Zhang committed
June 26, 2024
S
Update quantization README.md (#445)
supriyar committed
June 25, 2024
H
adding default inductor config settings (#423)
HDCharles committed
June 21, 2024
J
Refactor the API for quant method argument for quantize function (#400)
Jerry Zhang committed
H
077 autoquant gpt fast (#361)
HDCharles committed
June 17, 2024
J
Enable a test for loading state_dict with tensor subclasses (#389)
Jerry Zhang committed
June 13, 2024
J
Deprecate top level quantization APIs (#344)
Jerry Zhang committed