Introducing arXivQA: Training retrieval agents for arXiv search We curate a multi-hop arXiv dataset based on real queries and use rubric-as-rewards to train a Qwen model for production-grade retrieval Our latest blog details our experience training with both RLVR and rubrics
RESEARCH
-

Karpathy’s AutoResearch Enables AI Self-Improvement Loop
By
–
You can now give an AI a GPU and let it improve its own model overnight. This repo turns LLM research into an automated search loop. @karpathy just open-sourced a system that lets AI iterate on its own training code. Karpathy’s AutoResearch turns LLM improvement into an
-

NVIDIA CEO Jensen Huang on AI’s Five-Layer Stack Rebuild
By
–
The AI era is a full-stack rebuild, from power to models to real-world applications. Here’s NVIDIA CEO Jensen Huang on the five-layer AI cake driving the shift.
-

Self-Evolving AI Agents Tool Genesis Benchmark Research
By
–
New research on Self-Evolving AI Agents. Really interesting benchmark for evaluating a critical but overlooked capability: can LLMs create reusable tools from scratch, not just use existing ones? Tool-Genesis tests whether models can infer interfaces, generate schemas, and
-
Portfolio Strategy: Specialized Models Beat One-Size-Fits-All
By
–
The winning strategy is not one model to rule them all. It’s a portfolio: • reasoning models
• fast models
• research models
• coding agents
• multimodal systems
A model for every job. -
Space AI and Future Human Potential Convergence
By
–
Space, AI And The Future Of Human Potential As the space economy grows, the convergence of AI, medicine and space technology could unlock entirely new opportunities for humanity and business. Read more https://
bernardmarr.com/space-ai-and-t
he-future-of-human-potential/
… #SpaceTech #AI #Innovation #BernardMarr -

Health Dominates Copilot Mobile User Questions in 2025
By
–
In 2025, health was the #1 topic for Copilot mobile users. In our latest paper, we analyze over half a million conversations to understand what questions they're asking, and how to help.
-
New AI Results on Previously Unknown Technology Created
By
–
I'm getting strong results for technology the absolutely cannot be in the training data because I just created it myself
-

Structured-RAG: Solving Aggregative Queries in RAG Systems
By
–
Standard RAG falls apart on aggregative queries. Things like "average ARR for companies with >1k employees" require reasoning across hundreds of docs, and vector retrieval just can't handle that reliably. We built Structured-RAG to fix this:
– induces a schema from your -

Stanford study finds AI models degrade relationship advice
By
–
BREAKING: Stanford just proved that every major AI model is systematically making you worse at relationships. Researchers tested 11 state-of-the-art AI models, including GPT-5, GPT-4o, Claude, Gemini, Llama, DeepSeek, and Qwen, across thousands of real advice scenarios. The
