SIGN IN SIGN UP

[ET-VK] Store weights transposed for int8 linear (#9803)

## Context

The weight tensor of a linear layer is usually stored in a transposed manner, such that when computing the matrix multiplication, the reduction traverses along the rows of the weight tensor as opposed to the columns. This results in a better memory access pattern for CPUs.

However, for GPUs, I have found that "un-transposing" the weight tensors result in better performance. This is likely due to the fact since GPUs can compute multiple output elements in parallel, reading along the columns allows for coalescing memory loads among threads in a work group.

## Changes

* Introduce the ability to transpose height and weight dims when transferring tensor data to the GPU.
* Prepackthe weight tensor "un-transposed" for the int8 quantized linear operator

Differential Revision: [D72066588](https://our.internmc.facebook.com/intern/diff/D72066588/)
P
pytorchbot committed
655531ffca192dc44cfd578eafe770fc4f4f58aa
Parent: bcf4b46
Committed by GitHub <noreply@github.com> on 4/1/2025, 4:41:55 PM