Breaking: A free AI model trained for $7,800 has just outperformed a model 400 times larger in mathematics competitions. It is compact enough to run on a laptop. Weibo's AI lab published the results and put all the
AI
-

Low 20T token budget and bug halt DSV4 pre-training
By
–
Could it be linked to the (very) low 20T token pre-training budget? DSV4 was trained on 33T tokens. More room for knowledge in the case of HLE. Looks like they had to stop it early due to a bug they never root-caused. (Not a great demo for NVFP4 pre-training to be honest.)
-

Nemotron 3 Ultra unable to recover HLE and code performance via OPD
By
–
Nemotron 3 Ultra can't recover perf on HLE, code, etc. via OPD The teacher was trained on DeepSeek-V4-Pro traces (DSV4 Max achieves 37.7% on HLE!). Looks like the MOPD warmup failed to properly init the student? No good trajectory → No improvement via OPD
-
Why AI Literacy Is Now a Critical Boardroom and Investor Priority
By
–
Why AI Literacy Has Become A Boardroom And Investor Priority #AILiteracy is rapidly becoming a boardroom, #regulatory and #investor issue as companies move from AI experiments to real-world deployment. We explore why #businesses need measurable AI #education across the
-

TextPro-SLM Approach Bridges Speech and Text AI Gap
By
–
Why do speech AI models still lag behind text AI? Researchers at CUHK and Huawei propose TextPro-SLM — an approach that shrinks the gap by making spoken input look more like text input. Instead of tweaking the output, they redesign the input side with a unified speech
-

New directions in AI based on continual interaction and causality
By
–
In this blog, we explore new potential directions for the field of AI based on continual interaction and causality: https://love4all.ai/blog/continual-interactive-causal-agents/ … We have been working on this for years. Pedro Ortega pointed out the problem much earlier, when I
-
Subscription of $200 and loop without token spending in training
By
–
hehe luckily it has been with subscription ($200) and in a loop where there is a lot of execution without token spending during model training.
-

Learning Path for LLM Serving Engines: vLLM, SGLang, TensorRT-LLM
By
–
How to go about learning all of this? 1st: Start with the serving engine view – vLLM: PagedAttention, continuous batching, prefix caching, CUDA graphs – SGLang: RadixAttention/prefix reuse, speculative decoding, MoE, structured/agent workloads – TensorRT-LLM: NVIDIA peak
-

Confirming the growing popularity of edge models
By
–
Can confirm, edge models keep getting more and more popular
-
AI shifts from reactive programming to proactive orchestration
By
–
The deeper shift is from reactive programming to proactive orchestration, where AI anticipates needs rather than just responds. It's about designing systems that learn intent, not just behavior.
