@teortaxesTex About DeepSeek V4 being able to compete with the frontier, I made a new benchmark that suggests open models (particularly the new ultra-sparse ones) are qualitatively worse at problem solving and logical reasoning:
By
–

@teortaxesTex About DeepSeek V4 being able to compete with the frontier, I made a new benchmark that suggests open models (particularly the new ultra-sparse ones) are qualitatively worse at problem solving and logical reasoning: