Nick Karpov and Holly Smith walk through some of the latest Databricks features – and how they work together under a single architecture. Our R&D teams have been busy. Recent updates include:
– Three new foundation models in Databricks Foundation Model API support
– Stateless
RESEARCH
-

Databricks Releases New Foundation Models and Platform Updates
By
–
-

Market sentiment skepticism on intelligence capabilities
By
–
Same sentiment today on Intelligence. ‘stochastic parrot,’ ‘AI bubble’ ect.
-
Read more: alphaxiv.org/abs/2603.18886
By
–
read more:
alphaxiv.org/abs/2603.18886 [Translated from EN to English]→ View original post on X — @askalphaxiv, 2026-04-04 17:55 UTC
-

Principia: Mathematical Object Reasoning Benchmark for Frontier Models
By
–
“Reasoning over Mathematical Objects” Most reasoning benchmarks still let models answer with multiple choice or short numerics, which makes evaluation easy but also makes the task easier than real STEM reasoning. This paper shows that when you remove the options and ask for the actual object, like an equation, matrix, set, interval, or piecewise function, the performance drops sharply, even for frontier models. So this paper proposes Principia: a benchmark, training set, and verifier pipeline built specifically for mathematical-object reasoning, plus on-policy judge training to score these hard outputs reliably. What makes this interesting is that training on these harder outputs also improves standard math and science benchmarks, suggesting this is not just better formatting, but better actual reasoning.
→ View original post on X — @askalphaxiv, 2026-04-04 17:55 UTC
-
Nvidia Eyes Model Serving at 10,000-20,000 Tokens Per Second
By
–
Nvidia's Chief Scientist Bill Dally says there's a path to serving relatively large models at 10,000 to 20,000 tokens per user per second.
— Marcelo P. Lima (@MarceloLima) 4 avril 2026
For context, Opus 4.6 is ~43 and Grok 4.2 Beta is ~251 tokens/user/s 🤯 pic.twitter.com/mbZNFfWgUbNvidia's Chief Scientist Bill Dally says there's a path to serving relatively large models at 10,000 to 20,000 tokens per user per second. For context, Opus 4.6 is ~43 and Grok 4.2 Beta is ~251 tokens/user/s 🤯 [Translated from EN to English]
-
Hidden Gems: Underrated Datasets on Hugging Face Hub
By
–
Huggingface hub has so many underrated datasets ngl ! This is one among them, best part it can be used for both SFT and RL (if used properly) huggingface.co/datasets/jupy… Any other such underrated datasets/models/envs, drop them below
→ View original post on X — @clementdelangue, 2026-04-04 17:38 UTC
-
Human Cognitive Errors vs LLM Hallucinations Explained
By
–
So many people are confused about the relation between human cognitive errors and LLM hallucinations that I wrote this short explainer two years ago. Since many of those confusions persist, I am reposting:
-

RL rewards bias: next frontier is uncertainty tolerance
By
–
the direction became clear with AlphaGo 10 years ago, history repeats itself William Fedus (@LiamFedus) RL against verifiable rewards in LLMs has clearly opened a very powerful regime. It works, and because it works, there is a strong tendency to view more and more problems through that lens. You optimize for tasks where the reward is clean, where success is easy to check, where the feedback loop closes quickly. This is productive and will keep paying off. But it also creates a bias: you start emphasizing what is legible to the training setup, not necessarily what is most valuable. Scientific reasoning is a good example. Not every step in science is something that can be cleanly graded at the moment it is produced. A hypothesis can later fail experimentally and still have been exactly the right kind of thinking at the time: creative, mechanistically grounded, and responsive to the available evidence. “Turns out to be wrong” does not imply “was low-quality thinking”. A big part of the next frontier will be AI systems that can operate well under this kind of uncertainty, just like a big part of the last one was RL against verifiable rewards. — https://nitter.net/LiamFedus/status/2040462826851201256#m
→ View original post on X — @ceobillionaire, 2026-04-04 17:26 UTC
-

Complete Roadmap for Learning Agentic AI and Full-Stack Intelligence
By
–
Roadmap to learn Agentic AI 🚀 AI fundamentals Python + frameworks LLMs Agents architecture Memory + RAG Planning & decision-making RL & self-improvement Deployment Real-world automation Agentic AI = full-stack intelligence. Credit: Tiksly #AgenticAI #LLM #RAG #A
→ View original post on X — @ingliguori, 2026-04-04 17:25 UTC
-

Do Predictive Models Need to Be Causal?
By
–
Do predictive models need to be causal? – Biased and Inefficient buff.ly/tu50XsN #AI #MachineLearning #DeepLearning #LLMs #DataScience
→ View original post on X — @miketamir, 2026-04-04 16:40 UTC