Si se confirman estos resultados de Grok 4 -y según cómo se hayan calculado- estaríamos ante un modelo muy potente. Un 45% en Humanity's Last Exam es una salvajada. Más del doble de los modelos actuales (o3/Gemini)
LLMS
-

AI Models Achieve Genius-Level IQ for Business Scale
By
–
Some AI models now score in the genius IQ range—excelling at verbal reasoning, logic, and problem-solving. A powerful edge for businesses aiming to scale language-based analysis and decision-making. Source @VisualCap Link https://
buff.ly/swVkfqJ via @antgrasso #AI -
Grok 4 Leaked Benchmarks Show Significant Performance Gains
By
–
If the Grok 4 leaked benchmarks are right, it is going to be very useful that Humanity’s Last Exam has a holdout set of questions, because a rumored 45% score is a very big gain over the 20% or so of o3 & Gemini, and it would be pretty impressive (assuming no data contamination)
-
Unified AI Model Naming Could Reduce Consumer Confusion
By
–
a unified model would make things much clearer. when I talk to people who dont follow ai they almost always think o3 is an older version of 4o. usually means they see it fail at a complex problem not realizing it's not made for that.
-

Grok 4 Benchmarks vs Other Models
By
–


Grok 4 early benchmarks in comparison to other models. Humanity last exam diff is Visualised by @marczierer
-
Abacus.AI ChatLLM: All-in-One Platform for Emails, Projects, Images and Videos
By
–
All-in-One AI Platform for Everyday Tasks! Why settle for just chat when you can have it all? http://
Abacus.AI’s ChatLLM helps you draft emails, manage projects, generate images & videos, and even automate workflows.
No technical expertise needed—just your ideas and our -
USA AI breakthroughs: transformers, GPUs, diffusion models
By
–
happy birthday to the USA, the greatest country, and the origin of the following innovations: – Transformers
– Pre-training (web-scale next-token prediction)
– RLHF
– RLVR
– RL
– GPUs
– TPUs
– PyTorch
– word2vec
– reasoning models
– GANs
– diffusion models
– VLMs
– self-driving -

Grok 4 SOTA avec des scores élevés
By
–


BREAKING : Grok 4 will be SOTA – 35% on HLE, 45% with reasoning
– 87-88% on GPQA – 72-75% on SWE Bench (for Grok 4 Code) * Not official benchmarks, I plotted these based on the leaked scores and results from the web for other models -
GLM-4.1V-Thinking: Powerful Vision-Language Model for Multimodal Reasoning
By
–
GLM-4.1V-Thinking – a powerful new vision-language model for multimodal reasoning!
From STEM to GUI agents, it outperforms models 8x its size.
Open-source, scalable, and state-of-the-art.
Paper Link: https://
arxiv.org/abs/2507.01006 #AI #VLM #Multimodal #GLM4 #OpenSourceAI -
Share insights on LLMs, AI Agents, and Machine Learning
By
–
If you found it insightful, reshare with your network. Find me → @akshay_pachaar For more insights and tutorials on LLMs, AI Agents, and Machine Learning!
