Implement weight_int8packed_mm (#3893)
Summary: Pull Request resolved: https://github.com/pytorch/executorch/pull/3893 ## Context As title, this changesets implement the `aten._weight_int8packed_mm` operator. The operator implements a linear layer where the weight is quantized symmetrically to 8 bits for each "group". ghstack-source-id: 229569180 Reviewed By: copyrightly Differential Revision: D58263387 fbshipit-source-id: a32362ae912a4ad72bd597d8757528180fa3186b
S
Stephen Jia committed
a3ff00d3224633889b657e926ffd237324153e16
Parent: be7150c
Committed by Facebook GitHub Bot <facebook-github-bot@users.noreply.github.com>
on 6/10/2024, 8:17:06 PM