SIGN IN SIGN UP

[Executorch][llama] Change runner to decouple prompt length from sequencelength

Following previous diff now we can utilize entire kv cache to generate more
tokens than max prompt length allowed.

Differential Revision: D69073908
P
pytorchbot committed
dfa779618e57f3dfb5b538466cc83fd74bd0a957
Parent: 9fc101f
Committed by GitHub <noreply@github.com> on 3/26/2025, 11:24:13 PM