With all these new models dropping, that means… We are having a new release too! v2.2.6: Gpt 4.5, Claude Sonnet 3.7, Gemini 2.0 Flash Post Processing Arize & Phoenix Analytic Tavily Tool :
LLMS
-

GPT4.5 Vibe Test: Higher Emotional Intelligence and Creativity
By
–
Vibe Testing GPT4.5 OpenAI just released GPT4.5 w "higher emotional intelligence and creativity" than prior models. Here, we overview its capabilities and vibe test it on research / writing using our Open Deep Researcher. Video: https://
youtu.be/Jo2-fZ2THzU -
ChatLLM Research Tool: URL Scraping and Data Analysis Capabilities
By
–
Conduct your research with ChatLLM!
— Abacus.AI (@abacusai) 28 février 2025
It can scrape URLs, extract key insights, and perform analysis. While its depth is still evolving, we're constantly improving it for more in-depth research capabilities. pic.twitter.com/9MUU2txP93Conduct your research with ChatLLM! It can scrape URLs, extract key insights, and perform analysis. While its depth is still evolving, we're constantly improving it for more in-depth research capabilities.
-
GPU Constraints May Limit GPT-4.5 Rollout at OpenAI
By
–
You mean GPUs? Yeah, well, it depends. OpenAI is a much bigger company with many more regular users now. Maybe o1 and Deep Research users are hogging so many GPUs that they can't roll out GPT-4.5 to as many as they would have liked. Without more info it's hard to say anything
-

GPT-4.5 trained with new supervision techniques according to system card
By
–
Ok, according to the system card (
https://
cdn.openai.com/gpt-4-5-system
-card.pdf
…), it's trained "using new supervision techniques" -
RL and inference-compute scaling potential for model improvement
By
–
I wouldn't discount it quite, yet. I am curious how it'll do with some additional RL + inference-compute scaling à la o1.
-

Debugging agents with tracing and LLM-judge systems
By
–
How do I debug my agent? You can trace your agent run for later inspection, for instance using @ArizePhoenix
. @JohnGilhuly and team have just made a blog post explaining how to instrument a smolagent run, and how to setup LLM-judge systems. Should be mandatory reading for -

GPT 4.5 Generates Hilarious Content According to Dan Shipper
By
–
GPT 4.5 is actually so fucking funny > be me
> Dan Shipper -
Missing Technical Report Reduces Excitement About New Release
By
–
I agree, but at the same time I think people (or at least I) would have been much more excited about this release if it came with a Llama- or DeepSeek-style technical report.
It's hard to get excited about something if you a) don't know how it works or b) it's not substantially -
Claude 3.7 Sonnet Evaluated by US National Labs Scientists
By
–
We’re proud to participate in @ENERGY
’s first 1,000 Scientist AI Jam. Claude 3.7 Sonnet will be evaluated by National Labs scientists for research and national security applications, advancing US leadership and innovation via public-private partnership.