SIGN IN SIGN UP

Default Uninitialized Llama2 Weights to Zeros, and Provide Better Quantization for Example Models (#9634)

### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.

Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.

### Test plan
OSS XNNPACK CI for all model delegation


cc @digantdesai @cbilgin
M
Max Ren committed
91be93c4e14f5f11606ad0689ab0736bfd0aca6f
Parent: 46937eb
Committed by GitHub <noreply@github.com> on 3/26/2025, 7:59:11 PM