Feels like there's only 3 players left in the model game: Google, Anthropic, and Open AI.
LLMS
-
Perplexity Deep Research Dominates AI Research Tools Market
By
–
Perplexity Deep Research is the best in the industry right now
-
LLM Tool Call Vulnerability Discovered Across Major Providers
By
–
Piotr discovered something worrying: if you give an LLM a list of tools it's allowed to call, it might decide to also call a tool you didn't provide! Impacts all major US providers except @OpenAI
. Be sure to check LLM tool call requests! (Lisette/Claudette check automatically) -
Latest AI Models to Explore: Recommendations
By
–
What are the latest AI models to explore? any recs?
-
MaxRL: One-Line GRPO Improvement for Better Training Scaling
By
–
MaxRL is a slick one line change to GRPO that optimizes a maximum likelihood objective instead of expected reward
— alphaXiv (@askalphaxiv) 19 février 2026
Scaling advantages with 1/μ as opposed to traditional REINFORCE or GRPO’s 1/σ allows the model to make better progress from low initial pass rates and better… pic.twitter.com/sJCAlsezeqMaxRL is a slick one line change to GRPO that optimizes a maximum likelihood objective instead of expected reward Scaling advantages with 1/μ as opposed to traditional REINFORCE or GRPO’s 1/σ allows the model to make better progress from low initial pass rates and better
-

Automatic Prompt Caching API Feature Eliminates Manual Cache Points
By
–
Huge quality of life upgrade for devs: We've added automatic prompt caching to the API which means you no longer have to set cache points in your requests!
-
Gemini 3.1 Pro Doubles Reasoning Capability on ARC-AGI-2
By
–
Introducing Gemini 3.1 Pro 🚀
— Google AI (@GoogleAI) 19 février 2026
3.1 Pro represents a major step forward in core reasoning. It scored 77.1% (more than doubling 3 Pro’s score) on ARC-AGI-2, the benchmark that evaluates a model's ability to solve new logic patterns and work through challenges it hasn’t encountered… pic.twitter.com/8HOPoji7i5Introducing Gemini 3.1 Pro 3.1 Pro represents a major step forward in core reasoning. It scored 77.1% (more than doubling 3 Pro’s score) on ARC-AGI-2, the benchmark that evaluates a model's ability to solve new logic patterns and work through challenges it hasn’t encountered
-

Repeating Prompts Boosts Non-Reasoning Model Performance
By
–
now trending on alphaXiv when using non-reasoning model, simply repeating your prompt improves performance. basically 1-10% absolute improvement, depending on prompt ordering or the model. And this applies to every model, it wouldn't create extra latency or generated
-
Three Ways to Optimize LangSmith Agent Memory and Performance
By
–
LangSmith Agent Builder uses memory to improve with feedback. Three practical ways to get the most out of memory: → Tell your agent to remember what works
→ Use skills to give it specialized context when needed
→ Edit its instructions directly when that's faster Full

