SIGN IN SIGN UP

llama : default sampling changes + greedy update (#9897)

* llama : deprecate softmax sampler + fix dist sampler

ggml-ci

* tests : replace macros with functions

ggml-ci

* sampling : change temperature sampler logic

For t <= 0.0f, keep the max logit intact and set the rest to -inf

* cont : no need for special "greedy" logic

top-k == 1 is the same

* tests : init prob correctly

* llama : handle temp <= 0.0 in the temp_ext sampler too

ggml-ci

* cont : avoid extra loop in temperature sampler for sub-zero temp

ggml-ci
G
Georgi Gerganov committed
55e47786e373c90fc7803e718e3e1dd6d53c3db6
Parent: bc21975
Committed by GitHub <[email protected]> on 10/21/2024, 6:46:40 AM