META JUST KILLED TOKENIZATION !!!
@jiqizhixin
-
Search and Learn: Lightweight LLM Search Toolkit with vLLM
By
–
Search and Learn: A lightweight toolkit for implementing search strategies with LLMs and built for speed with vLLM. You can check it out here.
-
DVTS: Diverse Verifier Tree Search Improves AI Performance
By
–
Diverse Verifier Tree Search (DVTS): An unpublished extension we developed to the verifier-guided tree search technique. This simple yet effective method improves diversity and delivers better performance, particularly at large test-time compute budgets.
-
DeepMind’s Compute-Optimal Scaling for Open Models
By
–
this blog cover: Compute-optimal scaling: How we implemented DeepMind’s recipe to boost the mathematical capabilities of open models at test-time. https://
huggingface.co/spaces/Hugging
FaceH4/blogpost-scaling-test-time-compute
… -

Phi-4: Microsoft’s New Small Language Model for Complex Reasoning
By
–
Phi-4 is coming!!! Microsoft’s Newest Small Language Model Specializing in Complex Reasoning https://
techcommunity.microsoft.com/blog/aiplatfor
mblog/introducing-phi-4-microsofts-newest-small-language-model-specializing-in-comple/4357090
… -
Meta’s Coconut: LLMs Reasoning in Continuous Latent Space
By
–
Training LLMs to Reason in a Continuous Latent Space Meta presents Coconut (Chain of Continuous Thought), a novel paradigm that enables LLMs to reason in continuous latent space rather than natural language. https://
arxiv.org/abs/2412.06769 -

IEEE Announces 2025 Fellow Class Recognition
By
–
Congratulations!!! IEEE Fellow Class of 2025 announcement. https://
ieee.org/content/dam/ie
ee-org/ieee/web/org/about/fellows/fellow-committee/2025-fellows-class-announcement.pdf
… -

Densing Law of LLMs: Capability Density and Training Quality
By
–
Densing Law of LLMs https://
arxiv.org/pdf/2412.04315
v2
… introducing the concept of “capability density” to evaluate the training quality of large language models (LLMs) and describe the trend of LLMs that considers both effectiveness and efficiency. -
Diffusion Models Self-Guided Using Degraded Versions
By
–
Guiding a Diffusion Model with a Bad Version of Itself Tero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen, Timo Aila, Samuli Laine https://
arxiv.org/pdf/2406.02507 -
Not All Tokens Are Essential for Language Model Pretraining
By
–
Not All Tokens Are What You Need for Pretraining Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, yelong shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, Weizhu Chen https://
openreview.net/pdf?id=0NMzBwq
aAJ
…
