Making large AI models cheaper, faster and more accessible
[Inference] Optimize request handler of llama (#5512)
* optimize request_handler * fix ways of writing
傅
傅剑寒 committed
e6496dd37144202c8602dfdd66bb83f297eb5805
Parent: 6251d68
Committed by GitHub <noreply@github.com>
on 3/26/2024, 8:37:14 AM