SIGN IN SIGN UP

Update eager runner to support AttentionSink (#7149)

* Transform model to be able to use Attention Sink

Pull Request resolved: https://github.com/pytorch/executorch/pull/6700

This PR adds necessary functions for transforming the model to be able to use Attention Sink.
ghstack-source-id: 256108077
@exported-using-ghexport

Differential Revision: [D65571289](https://our.internmc.facebook.com/intern/diff/D65571289/)

* Update eager runner to support AttentionSink

Pull Request resolved: https://github.com/pytorch/executorch/pull/6921

This PR updates the eager runner to support AttentionSink.

It also fixes issues in the `chat_completion` function to properly handle the position id.
ghstack-source-id: 256108078

Differential Revision: [D66076486](https://our.internmc.facebook.com/intern/diff/D66076486/)

* add eval for attention sink (#7150)

Pull Request resolved: https://github.com/pytorch/executorch/pull/7070

This PR adds the function to evaluate the model's perplexity when AttentionSink is enabled.

This is mostly copied from https://github.com/mit-han-lab/streaming-llm/blob/main/examples/eval_long_ppl.py which is used by the AttentionSink paper to evaluate the model's perplexity when AttentionSink is enabled.
ghstack-source-id: 256108079
@exported-using-ghexport

Differential Revision: [D66474732](https://our.internmc.facebook.com/intern/diff/D66474732/)

Co-authored-by: Lunwen He <lwhecser@gmail.com>

---------

Co-authored-by: Lunwen He <lwhecser@gmail.com>
P
pytorchbot committed
5f0a14a2fa1c09e5439ee52b0a5135f33b6140f7
Parent: aa67cd9
Committed by GitHub <noreply@github.com> on 12/3/2024, 1:29:13 AM