Sonnet 4.5 aparece en la gráfica de METR colocándose en la línea de tendencia pero sin superar al punto más alto, ocupado por GPT-5, en este benchmark que mide el equivalente de horas de trabajo humano que un LLM puede resolver de forma autónoma.
LLMS
-

GPT-5 Pro Achieves Top Results on ARC-AGI Benchmarks
By
–
Un par de actualizaciones a benchmarks interesantes que estábamos esperando. GPT-5 Pro en ARC-AGI 1 y 2 logra ser el LLM disponible que mejores resultados consigue en ambos benchmarks.
-

Hugging Face Hub Launches Custom Apps, GGUF Editing, Xet Integration
By
–
The Hugging Face Hub team is on a tear recently: > You can create custom apps with domains on spaces
> Edit GGUF metadata on the Fly
> 100% of the Hub is powered by Xet – faster, efficient
> Responses API support for ALL Inference Providers
> MCP-UI support for HF MCP Server
> -

Microsoft delivers first NVIDIA GB300 cluster for OpenAI
By
–
Just announced: @Microsoft delivers the world's first at-scale NVIDIA GB300 NVL72 production cluster, providing the supercomputing engine needed for OpenAI to train multitrillion-parameter models in days, not weeks. Learn more: https://
nvda.ws/4ogROKu -
Small Fixed Data Can Poison Any Size LLM Model
By
–
Previous research suggested that attackers might need to poison a percentage of an AI model’s training data to produce a backdoor. Our results challenge this—we find that even a small, fixed number of documents can poison an LLM of any size. Read more:
-

Data-Poisoning Attacks Threaten LLMs Regardless Size
By
–
New research with the UK @AISecurityInst and the @turinginst
: We found that just a few malicious documents can produce vulnerabilities in an LLM—regardless of the size of the model or its training data. Data-poisoning attacks might be more practical than previously believed. -
GPT-5 Model Routing and Web Search Capabilities
By
–
Why? GPT-5 (auto) will likely route you to a GPT-5 model (there are many of them) that is less likely to do a web search and just produce answer your questions with results from the model itself. In the second prompt, you are asking it to access its own "past thinking" which the
-
GPT-5 Thinking vs Default: Hallucinations and Source Accuracy
By
–
AI can be confusing. How do you teach people that asking default GPT-5 a question and following up by asking for links to its sources will result in hallucinated cites while asking GPT-5 Thinking to answer a question and provide sources will get you accurate citations & links?
-

Deloitte AI Hallucinations Raise Oversight and Accountability Questions
By
–
Deloitte is refunding part of a $440K govt report after using AI (GPT-40), which hallucinated fake court cases + citations. Lawmakers called it a “human intelligence problem.”
AI didn’t fail, oversight did. What do you think? Image: theaifield | IG #AI

