feat: include tokens usage for streamed output (#4282)
Use pb.Reply instead of []byte with Reply.GetMessage() in llama grpc to get the proper usage data in reply streaming mode at the last [DONE] frame Co-authored-by: Ettore Di Giacinto <mudler@users.noreply.github.com>
M
mintyleaf committed
0d6c3a7d57101428aec4100d0f7bca765ee684a7
Parent: e001fad
Committed by GitHub <noreply@github.com>
on 11/28/2024, 1:47:56 PM