fp16 or bf16? I’m always a little nervous seeing people finetune or inference in fp16 models that were pretrained in bf16. The number of exponent bits (and hence range) is lower? I have a todo to look into it closer. Depends on the checkpoint possibly.
FP16 vs BF16: Precision Mismatch Concerns in Model Fine-tuning
By
–