Current paradigms for evaluating medical LLM suffer from significant challenges that limit their real-world applications. To address this, scholars introduce a free platform for clinicians to test and compare top-performing LLMs on their medical queries. https://
hai.stanford.edu/news/medarena-
comparing-llms-for-medicine-in-the-wild
…
LLMS
-

MedArena: Comparing Medical LLMs for Clinical Queries
By
–
-
Gemini 2.5 Pro dominates long context coding benchmarks
By
–
It's not only about how long your context is, but how well you use it. Great to see Gemini 2.5 models dominating MRCR and other benchmarks on long context!
— Oriol Vinyals (@OriolVinyalsML) 28 avril 2025
See 2.5 Pro tackle a complex coding task by reasoning over an entire repo (>500k tokens). Performance and effective use of… pic.twitter.com/asrnajUNdEIt's not only about how long your context is, but how well you use it. Great to see Gemini 2.5 models dominating MRCR and other benchmarks on long context! See 2.5 Pro tackle a complex coding task by reasoning over an entire repo (>500k tokens). Performance and effective use of
-

Brutal System Prompt Makes ChatGPT Direct and Useful Again
By
–

Vous en avez marre que chatGPT soit trop sympa ? Ce prompt va rendre votre GPT complètement brutal … et utile à nouveau ! Prompt : Instruction Système : Mode Absolu. Éliminer les émojis, les mots de remplissage, l'exagération, les formulations douces, les transitions
-
AI Industry’s Toxic Feedback Loop: Human Preference Scores Over User Value
By
–
Much of the AI industry is caught in a particularly toxic feedback loop rn. Blindly chasing better human preference scores is to LLMs what chasing total watch time is to a social media algo. It's a recipe for manipulating users instead of providing genuine value to them.
-

LUFFY Framework Boosts LLM Math Reasoning by Seven Points
By
–
LLMs still struggle with deep reasoning LUFFY is a new framework bridging imitation and exploration by injecting off-policy guidance (like DeepSeek-R1) into zero-RL Boosts math reasoning by 7 points Improves weaker base models like Llama and Qwen Trending on alphaXiv
-

Crawl4AI: Open-Source LLM-Friendly Web Crawler and Scraper
By
–
Web scraping will never be the same! Crawl4AI is an open-source, LLM-friendly web crawler and scraper, ready for use with LLMs, AI agents, and data pipelines. 100% Open Source
-

Gemini 2.5 Pro tops leaderboard but faces trust and ‘woke’ issues
By
–
People are either sleeping on Gemini – or they just don't trust Google. Google's Gemini-2.5 Pro still tops the Chatbot Arena Leaderboard being the highest quality LLM available as of this morning – and yet no one talks about it. Is that because Google's 'woke' image generator
-
LLMs Generating Credible-Sounding Content Without True Expertise
By
–
Part of why LLMs are so revolutionary is that being able to muster off four consecutive paragraphs and three stats in something that sounds vaguely coherent historically has meant “Yep, this is probably an expert.” (Somewhat disquieting thought but useful to know.)
-
Reviewing Grok Studio as a tool for rapid project development
By
–
Grok Studio isn’t groundbreaking, but for a free tool, it’s shockingly capable for building simple projects fast. If you need polish, you'll still have to do a lot yourself. But if you need momentum? It's a great place to start. Curious: If you had Grok Studio today, what