Partition Mutable Buffer as Core ML State (#5165)
* partition mutable buffer to coreml state * delegate llama mutable buffer to coreml * fix lint * support embedding quantize * try fix CI: 1. pin coremltools 8.0b2; 2. refrain from defaulting stateful llama until CI machine upgraded to MacOS 15 * address review comments: 1. add arg help info; 2. add mutable buffer partition log * fix CI: executorch example model test env is using older transformers, that does not support numpy 2.0 --------- Co-authored-by: yifan_shen3 <yifan_shen3@apple.com>
Y
Yifan Shen committed
f471556c05a26de383435cbf9f9896bb24f8ca0d
Parent: c5a385e
Committed by GitHub <noreply@github.com>
on 9/10/2024, 2:50:46 AM