llama : support batched embeddings (#5466)
* batched embedding: pool outputs by sequence id. updated embedding example * bring back non-causal attention * embd : minor improvements * llama : minor --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
D
Douglas Hanley committed
03bf161eb6dea6400ee49c6dc6b69bdcfa9fd3fc
Parent: ad014bb
Committed by GitHub <noreply@github.com>
on 2/13/2024, 12:06:58 PM