[Executorch][llama] Change runner to decouple prompt length from sequencelength
Following previous diff now we can utilize entire kv cache to generate more tokens than max prompt length allowed. Differential Revision: D69073908
P
pytorchbot committed
dfa779618e57f3dfb5b538466cc83fd74bd0a957
Parent: 9fc101f
Committed by GitHub <noreply@github.com>
on 3/26/2025, 11:24:13 PM