Update eager runner to support AttentionSink (#7149)
* Transform model to be able to use Attention Sink Pull Request resolved: https://github.com/pytorch/executorch/pull/6700 This PR adds necessary functions for transforming the model to be able to use Attention Sink. ghstack-source-id: 256108077 @exported-using-ghexport Differential Revision: [D65571289](https://our.internmc.facebook.com/intern/diff/D65571289/) * Update eager runner to support AttentionSink Pull Request resolved: https://github.com/pytorch/executorch/pull/6921 This PR updates the eager runner to support AttentionSink. It also fixes issues in the `chat_completion` function to properly handle the position id. ghstack-source-id: 256108078 Differential Revision: [D66076486](https://our.internmc.facebook.com/intern/diff/D66076486/) * add eval for attention sink (#7150) Pull Request resolved: https://github.com/pytorch/executorch/pull/7070 This PR adds the function to evaluate the model's perplexity when AttentionSink is enabled. This is mostly copied from https://github.com/mit-han-lab/streaming-llm/blob/main/examples/eval_long_ppl.py which is used by the AttentionSink paper to evaluate the model's perplexity when AttentionSink is enabled. ghstack-source-id: 256108079 @exported-using-ghexport Differential Revision: [D66474732](https://our.internmc.facebook.com/intern/diff/D66474732/) Co-authored-by: Lunwen He <lwhecser@gmail.com> --------- Co-authored-by: Lunwen He <lwhecser@gmail.com>
P
pytorchbot committed
5f0a14a2fa1c09e5439ee52b0a5135f33b6140f7
Parent: aa67cd9
Committed by GitHub <noreply@github.com>
on 12/3/2024, 1:29:13 AM