Fun to think about open-source models and their variants as families from an evolutionary biology standpoint and analyze "genetic similarity and mutation of traits over model families". These are the 2,500th, 250th, 50th and 25th largest families on @huggingface
:
LLMS
-

Open-source model families evolutionary analysis on Hugging Face
By
–
-
GPT-5 Model Gains Recognition Despite Controversy
By
–
GPT-5 is an amazing model. I can’t believe that’s controversial.
-

Uncertainty About GPT-5-Thinking Development Work and Improvements
By
–
we still dont know if they did extra work (in the last 4 months since o3/o4mini release) on the reasoner they call gpt-5-thinking or if they just packed up the existing o4 they had lying around internally and shipped it the very slight improvements in benchmarks suggests the
-

GPT 120B Performance Analysis: AWS and Azure Implementation Issues
By
–
Artificial analysis team did great work analysing the performance of the OpenAI's GPT OSS 120b model across providers – AWS and Azure are lagging behind. I really did not appreciate the extend to which the performance can drop if the model is not implemented correctly
-

GPT-5 Exceeds Medical Professionals on Reasoning Benchmarks
By
–
GPT-4o was below the level of medical professionals on medical reasoning benchmarks GPT-5 (apparently Thinking medium) now far exceeds them. (Usual benchmark caveats apply)
-
Weekly Roundup of Major AI Model Releases and Breakthroughs
By
–
What a crazy week in AI.. – OpenAI's GPT-5 Launch
– Anthropic's Claude Opus 4.1 Release
– Google's Genie 3 World Simulator
– ElevenLabs Music Generation Model
– xAI's Grok Imagine with 'Spicy' Mode
– Alibaba’s Qwen-Image Model
– Tesla AI Breakthroughs for Robotaxi FSD
– -
Prompt Engineering Impact on Test Outcomes Analysis
By
–
The size of these impacts is much larger than the effect of most prompt engineering on test outcomes.
-

Open Weight Model Performance Varies by Cloud Host Provider
By
–
This is actually a pretty surprising and something that should lead companies to change how they are thinking about hosting. Model performance for the open weights GPT model vary by meaningful amounts depending on who is hosting it, with Azure & AWS being low. Worth watching.
-
Genie 3 and Gemini 2.5 Advance Toward AGI with Better Benchmarks
By
–
To reach artificial general intelligence, we’ll need both advanced AI models that can think and understand the world around us and better benchmarks to evaluate their progress.
— Google AI (@GoogleAI) 12 août 2025
Listen in as @demishassabis and @OfficialLoganK chat about how our new world model Genie 3, Gemini 2.5… pic.twitter.com/aYSdmCqzyKTo reach artificial general intelligence, we’ll need both advanced AI models that can think and understand the world around us and better benchmarks to evaluate their progress. Listen in as @demishassabis and @OfficialLoganK chat about how our new world model Genie 3, Gemini 2.5
-
Comparing Kimi K2 and GPT-OSS Integration in Development Workflows
By
–
For those who’ve built with both Kimi K2 and gpt-oss, how have you integrated them into your workflows and which is working better for you?