Qualcomm AI Engine Direct - Support kv_cached stories 110M llama2 (#4142)
Summary: - Add custom memory descirptor - Add e2e example script verified with story110M in 8a8w, 16a4w - Add qnn_llama_runner to run static LLAMA. - Add readme - Add slice op test - Change RemoveClone to RemoveRedundancy - Change SimpleADB parameter artifact to build_path and related codes - Change multihead attentions to multiple single head. - Move sort inputs from execute to init - Remove split op - Support u16 and u8 mixed-precision quantization. Pull Request resolved: https://github.com/pytorch/executorch/pull/4142 Reviewed By: kirklandsign Differential Revision: D59339823 Pulled By: cccclai fbshipit-source-id: 51fcf14e406b04c51de6e421cccbad91a8ffa01e
S
shewu-quic committed
5584b9e3c865edca239ec5df6346f1d1aabb0276
Parent: 29fdaa1
Committed by Facebook GitHub Bot <facebook-github-bot@users.noreply.github.com>
on 7/6/2024, 12:04:23 AM