Are LLM leaderboards no longer trustworthy? @cohere
's deep dive reveals how Chatbot Arena scores can be gamed Providers test 10–27 private models & submit the best Proprietary models receive 2–3× more data Rankings may reflect overfitting Trending on alphaXiv
LLMS
-

LLM Leaderboards Gaming: Chatbot Arena Scoring Integrity Questioned
By
–
-
Gemini 2.5 Flash Preview Now Available on Poe
By
–
You can try Gemini 2.5 Flash at https://
poe.com/Gemini-2.5-Fla
sh-Preview
… and on all platforms. (2/2) -

Google Launches Gemini 2.5 Flash with Hybrid Reasoning
By
–
New: Gemini 2.5 Flash! Google DeepMind's latest model has hybrid reasoning capabilities, along with a 1m token context window, multimodal input support, and native search, all while prioritizing speed and cost. (1/2)
-

Loading Knowledge Base and Initializing RAG Agent for PDF Processing
By
–
5. Next, Loading the Knowledge & Agent When a PDF is processed, we instantiate PDFKnowledgeBase with the file path & our Milvus vector_db. Then we pass it into the above "get_rag_agent" method to initialize a fully ready-to-chat agent and then store it as a session variable.
-

Building a PDF RAG Agent with Agno and GPT-4o-mini
By
–
4. Next, let’s build the PDF RAG Agent The get_rag_agent function sets up the Agno Agent with gpt-4o-mini, PDFKnowledgeBase, and the DuckDuckGo tool for web search. This design allows the agent to try PDF knowledge base first and fall back to websearch if needed.
-

Meta Launches Autonomous AI App and AI News Roundup
By
–
Meta lance une application d'IA autonome
Mais aussi : les aperçus audio sont désormais multilingues dans NotebookLM, Figure AI bloque les ventes d'actions non approuvées et bien plus encore. Ma dernière newsletter gratuite : https://
vision-ia.beehiiv.com/p/meta-ia-auto
nome
… -

AdaR1: Adaptive Bi-Level Reasoning Framework for LLMs
By
–
Smart thinking ≠ long thinking.
Long chains of thought (CoT) help LLMs reason—but often, they’re overkill.
AdaR1, from a team of Chinese universities, introduces Bi-Level Adaptive Reasoning Optimization, a hybrid-CoT framework that dynamically mixes short & long reasoning chains -

The Future GUI of LLM Interfaces Will Be Visual, Not Text-Based
By
–
"Chatting" with LLM feels like using an 80s computer terminal. The GUI hasn't been invented, yet but imo some properties of it can start to be predicted. 1 it will be visual (like GUIs of the past) because vision (pictures, charts, animations, not so much reading) is the 10-lane
-
Understanding Large Language Models: Impact on Digital Transformation
By
–
New on the Blog: Understanding the Inner Workings of Large Language Models Stay ahead in the ever-evolving world of Digital Transformation. This article explores key insights into emerging technologies and their impact on today's business landscape. Discover how to
-
LLM Agent Generates Creative 3D Objects Part by Part
By
–
Built an LLM agent that constructs 3D objects part by part. A nice thing is you can leverage an LLM’s reasoning to generate out-of-distribution stuff like a chair with five legs. GPT-4o sometimes gets it; no luck with Imagen-3 yet.
— Yutaro Yamada (@_yutaroyamada) 1 mai 2025
Demoing today at #NAACL2025, Hall 3, 4-5:30pm! pic.twitter.com/g9Isda8q1xBuilt an LLM agent that constructs 3D objects part by part. A nice thing is you can leverage an LLM’s reasoning to generate out-of-distribution stuff like a chair with five legs. GPT-4o sometimes gets it; no luck with Imagen-3 yet.
Demoing today at #NAACL2025, Hall 3, 4-5:30pm!