Mistral AI released Voxtral TTS, a 3-billion-parameter text-to-speech model with open weights that the company says outperformed ElevenLabs Flash v2.5 in human preference tests roughly 63% of the time on standard voices and nearly 70% on voice customization. The model runs on
RESEARCH
-
OpenAI intern researcher release evaluation LLM research capabilities
By
–
I'm very curious about OpenAI's planned intern researcher release by September this year. Having tried using current LLMs for OpenAI's Golf Challenge, I would say that Codex & Opus are actively bad researchers (not meaningful difference between them)
– They come up with small -
New research paper shared via AlphaXiv
By
–
Paper: https://
alphaxiv.org/abs/2603.14473 Check out http://
AlphaSignal.ai to get a daily summary of top models, repos, and papers in AI. Read by 280,000+ devs. -

New OpenMOSS AI Framework Predicts Successful Science Experiments
By
–
Researchers trained an AI to predict which science ideas will succeed. Most AI helps scientists run experiments. This one decides which experiments are worth running. A new paper introduces OpenMOSS. This AI framework learned "scientific taste" from 700,000 pairs of
-

Human Traits AI Cannot Replicate and Their Value
By
–
The Human Traits #AI Can’t Replicate—And Why They’re Worth Developing
by @Forbes Learn more: https://
bit.ly/3PzSAWE #ArtificialIntelligence #MachineLearning #ML -

AI in Healthcare: Progress, Inequalities, Safety, and Real-World Deployment
By
–


Can #AI in healthcare really deliver – and for whom? Last night for #CamFest, we were joined by experts from across health & AI to ask what progress really looks like – from health inequalities to patient safety, regulation to real-world deployment. 💡
-

Discovery of a Structured Bug in Periodic Rollouts
By
–
2/4 The weird part: the spikes were periodic. When we increased rollouts per prompt, the spike pattern moved with them. That was the clue that this wasn't "training instability", it was a structured rollout-path bug. [Translated from EN to English]
-

Discovery of a logprob anomaly during Jamba training
By
–
1/4 We hit a strange logprob mismatch while training Jamba 3B with GRPO. Rollout logprobs and training-side recompute should match before any weight update. Ours didn't. That was the canary. 🧵 [Translated from EN to English]
-
ARC-AGI-4 Benchmark Release Planned Early 2027
By
–
For those wondering about ARC-AGI-4 timing: it will be released in early 2027. We are aiming for a yearly release schedule for new benchmarks. We are also aiming for each new benchmark to be fully unsaturated upon release, and to target the most important unanswered research
-
AI21Labs découvre débordement uint32 dans kernel CUDA vLLM Mamba-1
By
–
Thanks to @AI21Labs for tracking down a silent uint32 overflow in vLLM's Mamba-1 CUDA kernel and contributing the fix. Root cause: `uint32_t` stride × cache_index overflows silently at scale. Fix merged in #35275. The debugging story is worth a read. 🔗 ai21.com/blog/vllm-cuda-inte…
