𝕀𝔸 : 𝕝'𝕖𝕞𝕡𝕚𝕣𝕖 𝔸𝕞𝕒𝕫𝕠𝕟 𝕔𝕠𝕟𝕥𝕣𝕖-𝕒𝕥𝕥𝕒𝕢𝕦𝕖 Amazon dévoile Rainier son supercalculateur Nova, les 6 cavaliers du numérique pour contrer OpenAI L’arrivée de la 3D dans l’IA avec World Labs et Google Un clone pour Noël ? Et bien plus encore
LLMS
-
Major VLM Revolution: Open-Source Models from Google, OpenGVLabs, Qwen, Microsoft
By
–
VLMs are going through quite an open revolution AND on-device friendly sizes: > Google DeepMind w/ PaliGemma2 – 3B, 10B & 28B > OpenGVLabs w/ InternVL 2.5 – 1B, 2B, 4B, 8B, 26B, 38B & 78B > Qwen w/ Qwen 2 VL – 2B, 7B & 72B > Microsoft w/ FlorenceVL – 3B & 8B (Links below)
-
Llama 4 in Development: Community Celebrates Naming Choice
By
–
Yes! But I’m much more happy that they didn’t call it 3.1-Instruct-New Also, Llama 4 cooking!
-
O1 model reception shifts rapidly among users
By
–
fun watching the vibes shift so quickly on o1 🙂 glad you like it!
-

Phrase Matching Techniques in Marginalia Search Engine
By
–
Phrase Matching in Marginalia Search https://
bit.ly/4gf76v4 #AI #MachineLearning #DeepLearning #LLMs #DataScience -
SFT and RL Fine-Tuning: Complementary Approaches for Model Optimization
By
–
I think that SFT will remain useful for fine-tuning a model to a new specific task, and RL fine-tuning becomes an interesting additional toolset further to push the model toward a desired kind of answer while also allowing it more flexibility in the way it “thinks” of the answer
-
OpenAI’s RLFT Boosts o1-mini Accuracy in Law, Finance, Engineering
By
–
OpenAI sees huge potential in domains like Law, Finance, and Engineering, where “objectively correct” answers matter. Early results are promising: RLFT helped o1-mini boost accuracy in a gene example from 25% to 31%!
-
Question-Answer Pairs Training Without Perfect Ground Truth
By
–
Why is this interesting? It lets us use question-answer pairs without needing the full chain-of-thought input or even a perfect ground-truth output for SFT. Graders can assign nuanced scores (0–1), improving model outputs incrementally—even when answers are partially
-
RLFT Evaluates Complete Answers Instead of Individual Tokens
By
–
Instead of evaluating each token, one at a time, RLFT grades entire answers—perfect for scenarios like OpenAI’s o1, where we can’t access intermediate tokens before the “end_of_thought.”
-
OpenAI Releases Reinforcement Fine-Tuning for Flexible Model Training
By
–
Today, we got Reinforcement Fine-Tuning (RLFT)! (OpenAI releases day 2) Unlike Supervised Fine-Tuning (SFT) (already available on OpenAI for models like 4o), RLFT trains models with a flexible, non-fixed objective.
