Microsoft was on that – they call GPTS "agents"
LLMS
-

Cat facts in prompts derail AI reasoning, error rate up 10x
By
–

Apparently you can completely derail AI's ability to reason when you throw in "cat facts" or other out-of-context materials into the prompt. Error rate goes up 10x. When you are facing down a Terminator, just tell it that cats sleep for most of their lives.
-
Hybrid o4 Model with Agentic Capabilities and Web Navigation
By
–
Si me preguntáis a mí, también voy en la línea de un modelo híbrido (que sería o4) con mejoras en razonamiento, usos de herramientas y capacidades agénticas integrado en un ChatGPT vitaminado más parecido a lo que hace Manus, incluyendo Operator para navegar la web visualmente,
-

Jamba Open Model Family Gets New Update with Improved Grounding
By
–
Now live. A new update to our Jamba open model family Same hybrid SSM-Transformer architecture, 256K context window, efficiency gains & open weights. Now with improved grounding & instruction following. Try it on AI21 Studio or download from @huggingface More on what
-
What Should GPT-5 Deliver to Truly Surprise Users?
By
–
Estamos a días de que GPT-5 sea una realidad! Os pregunto… ¿Qué debería presentar OpenAI para que realmente GPT-5 os sorprendiera frente al combo de modelos que ya tenemos actualmente?
-
Future test may struggle with ungrammatical or misspelled text.
By
–
Will try in a future test. I suspect it would have a much harder time with that, especially for OOD (e.g. ungrammatical or misspelled) text.
-

ChatGPT o3-pro reads 1965 I.J. Good quote from torn strips
By
–

ChatGPT o3-pro identifies a 1965 quote by I. J. Good hand-written in a mix of print and cursive on a note ripped into four strips in reverse order rotated 90° in alternating directions:
-

Google Gemini 2.5 Flash-Lite: Fastest New AI Model Update
By
–
Google just dropped another monster AI update in June — and it’s PACKED! From genome-level breakthroughs to voice-first AI search to generative wallpapers from Jupiter … Here’s what’s new in AI @Google : Gemini 2.5 Flash-Lite: their fastest + most efficient model
-

Multimodal LLMs Fail 3D Puzzles: MARBLE Benchmark
By
–
Multimodal LLMs can write essays.
They can chat, caption, and even summarize papers. But give them a Portal map or a 3D puzzle, they break instantly. Zero percent accuracy. Welcome to the edge of AI reasoning: The MARBLE Benchmark