Polymarket has 97% odds on OpenAI releasing a new frontier model before Dec 31st, but only 16% that it's the best model by the end of the year? (Every GPT announcement has held the #1 position upon release)
LLMS
-

Jais-2: Arabic LLM 20x Faster Than Leading Models
By
–
Introducing Jais-2 – a major leap for Arabic AI. Co-developed with @G42ai , Inception, and @mbzuai
, Jais-2 is a family of Arabic models built end-to-end on Cerebras from pretraining to inference. Jais-2 70B runs at 2,000 tokens/sec – 20x faster than leading LLMs. It brings deep -
Meta’s Avocado AI Model Delayed to 2026
By
–
Meta is working on a new proprietary frontier AI model called “Avocado” according to CNBC. “We were expecting the model to be released before the end of this year, but that the plan now is for that to happen in the first quarter of 2026.” Meta will be back
-

Waymo’s Use of Large Models and GOFE in Autonomous Driving
By
–
My sense is that @Waymo has been using large models for awhile (plus lots of GOFE for sensor pre-processing, system bootstrapping, data avalanching, and ongoing safety). One clue is the “thinking fast and slow” note on the image.
-

RL Impact on Base Model Performance: Pre-training and Mid-training Interplay
By
–
There are competing views on whether RL can genuinely improve base model's performance (e.g., pass@128). The answer is both yes and no, largely depending on the interplay between pre-training, mid-training, and RL. We trained a few hundreds of GPT-2 scale LMs on synthetic GSM-like reasoning data from scratch. Here are what we found: 🧵
-
AI model coding performance on SWE benchmark
By
–
ca veut dire que c'est un model qui code très bien . SWE c'est le benchmark, 68 le score 🙂
-
Top AI Models for Different Tasks: A Practical Guide
By
–
ChatLLM model picks we keep coming back to: Everyday → GPT- 5.1 Coding → Sonnet 4.5 / Opus 4.5
Images → Nano Banana Pro, Midjourney → Sora 2, Seedance Pro
TTS → ElevenLabs, Hume
Fast tasks → Gemini Flash
Reasoning → GPT-5.1 Thinking Model switching handled -

Offline AI Coding Model Achieving 68 SWE Performance
By
–
MAIS WWWWWWWWWWHAAAAAAT !!! Mais purée, si c’est vrai, vous n’imaginez même pas le BANGER ASTRONOMIQUE !! Genre pouvoir utiliser une IA qui code à 68 SWE en OFFLINE !!!! Dans l’avion, ou quand t’as pas de réseau dans le train… Je vais tester ça tout de suite, mais ma hype est
-

Notion Testing GPT-5.2 Internally
By
–
BREAKING : Notion might be testing GPT-5.2 internally under the "olive-oil-cake" codename. The codename is abstracted by Notion. Soon or not soon?
-

Fireworks AI Achieves Top Performance with Kimi K2 Blackwell
By
–
Fireworks AI hits top performance on the Artificial Analysis leaderboard with Kimi K2 on NVIDIA GB200 NVL72 systems. @FireworksAI_HQ CEO Lin Qiao explains why Blackwell NVL72 transforms massive MoE serving. Learn more https://
nvda.ws/48rJ4Mm
