[Excutorch][Llama] Decouple input sequence length from kv cache context length (#8047)
Pull Request resolved: https://github.com/pytorch/executorch/pull/7927 Decouple max sequence length, for shape dynamism in torch.export, from sequence length used for kv cache sizing. ghstack-source-id: 263653316 Differential Revision: [D68448334](https://our.internmc.facebook.com/intern/diff/D68448334/) Co-authored-by: Kimish Patel <kimishpatel@fb.com>
P
pytorchbot committed
afc5a50c86b2a6af425d9d8868d623f6b69eb35e
Parent: 21ec3a1
Committed by GitHub <noreply@github.com>
on 1/30/2025, 8:53:11 PM