

Claude Opus 4.6 claimed the first spot on Arena across text, code and expert categories. Opus 4.6 scored 1496 on text

By
–


Claude Opus 4.6 claimed the first spot on Arena across text, code and expert categories. Opus 4.6 scored 1496 on text
By
–
Anthropic keeps working on Knowledge Bases, as a new "Save to knowledge base" button has been spotted in testing. Isn't this a continuous learning solution?
— 🚨 AI News | TestingCatalog (@testingcatalog) 5 février 2026
Save button triggers this prompt 👀
"Review this entire conversation and save any important, reusable information to the… https://t.co/04PuKmNqSM pic.twitter.com/xFirhgbqdB
Anthropic keeps working on Knowledge Bases, as a new "Save to knowledge base" button has been spotted in testing. Isn't this a continuous learning solution? Save button triggers this prompt "Review this entire conversation and save any important, reusable information to the

By
–

Claude Opus 4.6 is a new SOTA model on ARC-AGI-2 benchmark with 68.8% achievement. The next leap

By
–
OpenAI opens up Trusted Access framework to accelerate cyber defence. GPT-5.3-Codex was the first model to hit a "High" on OpenAI's preparedness framework. Shit is about to get real

By
–
BREAKING : GPT‑5.3‑CODEX WAS USED TO SUPPORT CREATING ITSELF, ACCORDING TO OPENAI'S BLOG! It achieves SOTA score of 57% at SWE Bench Pro and 76% on TerminalBench. "With GPT‑5.3-Codex, Codex goes from an agent that can write and review code to an agent that can do nearly

By
–
SOTA ON SWE BENCH "GPT‑5.3-Codex achieves state-of-the-art performance on SWE-Bench Pro"

By
–

Opus 4.6 comes with a big improvement at Agentic Search, Agentic financial analysis and Office tasks. "Financial professionals use AI to research across multiple data sources, support financial analyses, and create deliverables that their teams and customers can act on."
By
–
BREAKING 🚨: Claude Opus 4.6 has been officially announced. Opus 4.6 comes with an improved performance across various agentic, reasearch and coding tasks.
— 🚨 AI News | TestingCatalog (@testingcatalog) 5 février 2026
What would you test first? 👀 https://t.co/jZnbLKuMlx pic.twitter.com/dwpxCSfQSY
BREAKING : Claude Opus 4.6 has been officially announced. Opus 4.6 comes with an improved performance across various agentic, reasearch and coding tasks. What would you test first?
By
–
BREAKING 🚨: @perplexity_ai LAUNCHES MODEL COUNCIL, A NEW MODE WHERE GEMINI 3 PRO, OPUS 4.5 AND GPT 5.2 WILL WORK AS A SWARM OF ASYNC AGENTS ON YOUR TASK.
— 🚨 AI News | TestingCatalog (@testingcatalog) 5 février 2026
Perplexity MAX 🔥 https://t.co/n0eKrkhSw1 pic.twitter.com/sCvVQQwGsf
BREAKING : @perplexity_ai LAUNCHES MODEL COUNCIL, A NEW MODE WHERE GEMINI 3 PRO, OPUS 4.5 AND GPT 5.2 WILL WORK AS A SWARM OF ASYNC AGENTS ON YOUR TASK. Perplexity MAX
By
–
BREAKING 🚨: OPENAI ANNOUNCED OPENAI FRONTIER, A NEW ENTERPRISE PLATFORM TO CREATE AND MANAGE AI COWORKERS.
— 🚨 AI News | TestingCatalog (@testingcatalog) 5 février 2026
IT IS HAPPENING 👀
"Frontier gives agents the same skills people need to succeed at work: Understand how work gets done, Use a computer and tools, Improve quality over… https://t.co/5TyCHQ2EpP pic.twitter.com/wGbvqrxmXe
BREAKING : OPENAI ANNOUNCED OPENAI FRONTIER, A NEW ENTERPRISE PLATFORM TO CREATE AND MANAGE AI COWORKERS. IT IS HAPPENING "Frontier gives agents the same skills people need to succeed at work: Understand how work gets done, Use a computer and tools, Improve quality over