Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
Fix zero-grad behavior when entering the validation loop (#18710)
Co-authored-by: Jirka Borovec <6035284+Borda@users.noreply.github.com>
A
Adrian Wälchli committed
a26424e89ef71ff6cbfd8f1c89267df81a3184eb
Parent: 7fd5c02
Committed by GitHub <noreply@github.com>
on 10/9/2023, 10:07:54 PM