Stellar thread on the capabilities of o4 and some of the backfilling and hallucinations. Great benchmark
LLMS
-

Azure DeepSeek R1 Integration Tutorial with LangChain Package
By
–
Azure + DeepSeek Tutorial Learn to implement the DeepSeek R1 reasoning model with our new langchain-azure package. Simplified authentication and integration lets you build advanced AI applications faster. Watch the tutorial: https://
youtube.com/watch?v=aBkpzU
Bg6qs
… -
o1 Pro outperforms o3 for business and writing tasks
By
–
Not a coder, but for my tasks (business strategy, writing, etc) I'm still loving o1 Pro vs o3. o3 definitely has some special magic that o1 Pro doesn't. I'm with you on my interest in o3 Pro. Hope we see it soon.
-

OpenAI o3, Gemini 2.5 Pro, Claude 3.7 top AI IQ rankings
By
–
OpenAI o3, Gemini 2.5 Pro, and Claude 3.7 Sonnet Extended top the IQ charts among AI models, per Tracking AI. For reference, average human IQ = 100.
-
Why o3 crushes tasks: larger context, lower latency, better instruction, safety
By
–
Why o3 crushes these tasks • Larger context window than Sonnet 3.7 → fewer truncations
• Lower latency than Gemini 2.5 Pro → real‑time iteration
• Stronger instruction following than Grok 3 → cleaner outputs
• Fine‑tuned safety layer lets you push edge cases without -

10 Ways to Use ChatGPT-o3 Like a Pro as Top AI Model
By
–
I tested Gemini 2.5 Pro, Claude Sonnet 3.7, and Grok 3… But they are nowhere close to ChatGPT o3 and o4 mini. Here are 10 ways to use ChatGPT-o3 like a pro and unlock its full potential:
-
Call for OpenAI to fix hallucination issues in models
By
–
OpenAI should make some changes in the new models for hallucination mishaps. it does hallucinates and it sucks.
-
Switching Between Gemini and Sonnet in WindSurf Development
By
–
I switch back and forth in WindSurf often. Sometimes when Gemini can't solve it, Sonnet can (and vice versa).
-

Ollama Works Great for AI Development
By
–
Working great! We love Ollama.❤️ https://t.co/EDSs3oNDYy pic.twitter.com/LkfHnwdADF
— 机器之心 JIQIZHIXIN (@jiqizhixin) 19 avril 2025Working great! We love Ollama.
-
1B 4B 12B Models Running on Consumer GPUs
By
–
The 1B, 4B, and 12B parameter version should run on many consumer GPUs.