The evals that folks care about and publish at any time are those that SoTA models are just shy of being good at. So this is what such pics *always* look like.
MARKET TRENDS
-
Gradual Tech Improvement Expected Until Field Maturity
By
–
Ongoing gradual minor improvement in product capabilities is what you'd expect in pretty much any tech field at pretty much any time, until that field is extremely mature and stable.
-

Perplexity Enterprise Adoption in Legal Services Industry
By
–
An important vertical that Perplexity is working on for enterprise adoption is legal services. Lawyers want the most accurate AI that can pull references reliably. One of the best law firms Gunderson Dettmer has adopted Perplexity Enteprise and the results have even outstanding.
-

GPT-5.2 Demonstrates Rapid Progress in AI Frontier
By
–
GPT-5.2 shows how fast the frontier is moving. Impressive progress in just a few weeks.
-

OpenAI Launches GPT-5.2 Across Multiple Subscription Plans
By
–
OpenAI has officially launched GPT-5.2 Rolling it out across Plus, Pro, Business, and Enterprise plans, with Free and Go users gaining access tomorrow. The new system comes in tiers
– Instant
– Thinking
– Pro Each aimed at different levels of speed, reasoning, and accuracy. -

GPT-5.2 Now Available for Perplexity Pro and Max
By
–
GPT-5.2 is now available for all Perplexity Pro and Max subscribers.
-

GPT-5.2 Outperforms Gemini and Claude
By
–
GPT-5.2 is here and it just outsmart Gemini and Claude in Thinking evals Benchmarks Here's everything you need to know: 1/ GPT-5.2 basically ate the entire knowledge economy. It beats or ties industry experts on 70.9% of real-world knowledge work across 44 occupations:
-

GPT-5.2 Rollout and SOTA Performance
By
–
BREAKING : GPT-5.2 Instant, Thinking, and Pro are rolling out on ChatGPT. Free and Go users will receive it a bit later. Also, GPT-5.2 Pro (High) is SOTA for ARC-AGI-2, scoring 54.2% for $15.72/task! Big at coding and math
-

GPT-5.2 stats exclude Grok: Is there beef between companies?
By
–
GPT-5.2 comparison stats exclude Grok… I wonder if the two companies have a beef or something
-

Rerank 4 Outperforms Models for Enterprise Search Tasks
By
–
Rerank 4 outperforms other models for enterprise use cases across key domains, including finance, healthcare, and manufacturing search tasks.