GPT-120B is getting exceptional scores on maths benchmarks, often well above models like o3, Gemini 2.5 Pro, DeepSeek R1, and most others. What’s also remarkable is that it achieves this at a cost of $1-2 per benchmark, often 10-20x cheaper than equivalent proprietary models
LLMS
-
LLMs capabilities: high and low packed closely together
By
–
LLMs’ capabilities are like porcupine quills: very high and very low packed closely together.
-
Enterprise AI Agents Shift Focus to Behavioral Evaluation
By
–
As demand for smarter, more nuanced agents grows, enterprises are shifting their focus to evaluation. Databricks VP of AI @NaveenGRao and Chief AI Scientist @jefrankle explain why focusing on behavioral evaluation over model size is important and how innovations like Test-Time
-
MIT Benchmark Measures Chatbots’ Emotional Intelligence and User Behavior
By
–
The GPT-5 backlash highlights the fact that chatbots have little social or emotional intelligence. This week's AI Lab looks at a benchmark from MIT researchers that would attempt to gauge a model's capacity to encourages healthy behavior in its users.
-

Anycoder Model List Growing with Mobile Vibe Coding Support
By
–
Anycoder model list will continue to get bigger and you can use it for vibe coding on your phone
-

Balancing Empathy and Sycophancy in Conversational AI Systems
By
–
I feel like @parmy & @frimelle are introducing a very important nuance in the sycophancy debate with this article. I'm sure this is very challenging, especially at chatgpt scale but in my opinion, there's a way to keep conversational AI warm/empathetic/somehow sycophantic while
-

GPT-5 Thinking differs between Plus and Pro tiers
By
–
GPT-5 thinking in plus tier is not the same as GTP-5 thinking in pro tier. That „thinking“ (reasoning) setting is higher. https://
x.com/scaling01/stat
/scaling01/status/1955610515134460285
… -
Prompts as the key to unlocking AI chatbot potential
By
–
this is spot on, prompts are the key to unlocking chatbot potential
-
Prompt Engineering as a Form of User Experience Design
By
–
Prompt Writing is UX You’re not talking to a robot.
You’re designing how the AI thinks. Prompting is a language. Master it, and you control the conversation. -
Techniques to Reduce AI Model Hallucinations
By
–
Bonus: Reduce Hallucination Hallucinations happen when models make stuff up. Fix it with: Retrieval Augmented Generation (RAG) – ReAct (reason + act)
– Chain-of-Verification Don’t just ask questions. Ask it to check its own answers.
