llama : add qwen2moe (#6074)
* support qwen2moe * fix-review * metal : support unary ops for nelements % 4 != 0 * metal : require contiguousness for float4 unary kernels * metal : require contiguousness for float4 unary kernels (cont) * fix-review * names : for brevity "SHARED_EXP" -> "SHEXP" * llama : reuse build_moe_ffn() * llama : add model type name --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
S
Shijie committed
f4dea7da1841a92d2788b0535063abf2f0e28461
Parent: 8a56075
Committed by GitHub <noreply@github.com>
on 4/16/2024, 3:40:48 PM