Yes I agree it is, but that's what makes it bad, it doesn't mean that there's no human who could use them for research, same like you could probably build decent software with Sonnet 3.5 – but it was still a bad model vs what we have now
GENERATIVE AI
-

Mistral AI Launches Voxtral TTS Model Outperforming ElevenLabs
By
–
Mistral AI released Voxtral TTS, a 3-billion-parameter text-to-speech model with open weights that the company says outperformed ElevenLabs Flash v2.5 in human preference tests roughly 63% of the time on standard voices and nearly 70% on voice customization. The model runs on
-
AI code prediction: hype or solid reality, adoption is key
By
–
Two notes from this year-old prediction:
— Ethan Mollick (@emollick) 26 mars 2026
1) You can either view this as hype (100% of code is not written by AI) or a startlingly solid prediction (Claude Code didn’t exist then, but now writes a remarkably high percentage of code)
2) Adoption is more of a barrier than technology https://t.co/z7kG2be8hfTwo notes from this year-old prediction:
1) You can either view this as hype (100% of code is not written by AI) or a startlingly solid prediction (Claude Code didn’t exist then, but now writes a remarkably high percentage of code)
2) Adoption is more of a barrier than technology -
OpenAI intern researcher release evaluation LLM research capabilities
By
–
I'm very curious about OpenAI's planned intern researcher release by September this year. Having tried using current LLMs for OpenAI's Golf Challenge, I would say that Codex & Opus are actively bad researchers (not meaningful difference between them)
– They come up with small -

OpenAI Shelves Adult Chatbot Plans Over Safety Concerns
By
–
OpenAI has indefinitely shelved its planned "adult mode" erotic chatbot amid pushback from staff and investors over risks to minors and concerns about encouraging unhealthy emotional attachments to AI. The decision is part of a broader refocusing away from "side quests" toward
-

Fine-tuning Gemini 2.5 made it worse
By
–
HOLY SHIT… Google AI just proved that fine-tuning Gemini 2.5 made it dumber on hard queries. > Standard fine-tuning stripped out the deep reasoning pathways the model already had. Replaced them with shallow pattern matching. The fine-tuned version scored lower than the base
-

Top AI Stories: AGI Testing, Bot Moderation, GIF Tools
By
–
Top stories in AI today: – ARC’s new AGI test stumps every frontier AI
– Reddit's AI bot crackdown skips the ID check
– Create branded reaction GIFs for Slack
– Google shrinks AI memory with zero accuracy loss
– 4 new AI tools, community workflows, and more -
GPT 5.4 Pro API Access and Pricing Considerations
By
–
Good request! I have an internal skill for this which allows you to use GPT 5.4 Pro via the API and provide it the right context but I’m not sure if this makes as much sense price wise!
-

Nemotron: NVIDIA’s Underrated AI Model Breakthrough
By
–
Nemotron remains completely underrepresented. NVIDIA has cooked up a storm, and most people are still ignoring it.
-
Genie 3: From Content Generation to Interactive Reality
By
–
Genie 3 signals where frontier AI is going next: From generating content → to generating interactive realities. That changes how we think about R&D, simulation, and even product design. I break it all down in this latest video. Don't miss out on the latest AI advancements!