Could it be linked to the (very) low 20T token pre-training budget? DSV4 was trained on 33T tokens. More room for knowledge in the case of HLE. Looks like they had to stop it early due to a bug they never root-caused. (Not a great demo for NVFP4 pre-training to be honest.)
Low 20T token budget and bug halt DSV4 pre-training
By
–
