
GPT-5 Benchmarks 42% on HLE and 75% on SWE Bench
By
–
GPT-5 mini on o3 level, when free user hit rate limit

By
–
GPT-5 just dropped and the benchmarks look insane It is incredible at single-shot understanding of complex prompt, multimodal reasoning and writing code that just works. Huge leap from previous generation of OpenAI, Claude and Qwen AI models.
By
–
Get started by reading our blog post: https://
blog.langchain.com/introducing-al
ign-evals/
…
Or, watch a video on how it works:

By
–
Want more consistent, reliable evaluation scores? Last week we launched Align Evals in LangSmith — helping you craft better prompts to reduce inconsistencies, and build more reliable evals More details in the reply below
By
–
OpenAI secretly gave me access to GPT-5 for weeks and I've tested it thoroughly for you. Here is everything it can do:
By
–
GPT-5 is a very good writer, able to connect analogies between things you wouldn't expect older GPT series models to do. There are actual insights now in what it produces over generic word salad. Deep research also started returning niche numbers/insights that I didn't expect. 5