o3 is very performant. More importantly, progress from o1 to o3 was only three months, which shows how fast progress will be in the new paradigm of RL on chain of thought to scale inference compute. Way faster than pretraining paradigm of new model every 1-2 years
LLMS
-

OpenAI o3 Breakthrough Reasoning Model Enters Safety Testing
By
–
o3, our latest reasoning model, is a breakthrough, with a step function improvement on our hardest benchmarks. we are starting safety testing & red teaming now.
-

Rapid Progression of AI Models: Toward Superintelligence by 2030
By
–
September 2024: o1-preview had an IQ of 120 December 12, 2024: o1 displayed an IQ of 133 o3, successor to o1, just gained 15 points in mathematics. It becomes increasingly plausible that AI could be a million times smarter than humans by 2030
-
OpenAI Launches o3-mini and o3 Models for Safety Testing
By
–
if you are a safety researcher, please consider applying to help test o3-mini and o3. excited to get these out for general availability soon. extremely proud of all of openai for the work and ingenuity that went into creating these models; they are great.
-

OpenAI reveals o3 and o3-mini models
By
–

BREAKING : OpenAI revealed o3 and o3-mini. o3 is a new model which is 20% better at programming than o1 o3 mini is planned to be launched publicly by the end of January
-
OpenAI Launches o3 and o3 Mini for Safety Testing
By
–
OpenAI announcing o3 and o3 mini today! Not launching, but allowing public access for safety testing. It scores 71% on SWE BENCH—over 20% better than o1 Coding is going to be changed forever
-
Generate SQL Commands with Natural Language Using ChatLLM Teams
By
–
You can use natural language to generate your #SQL commands using @AbacusAI ChatLLM Teams! Get FREE TRIAL at https://t.co/73VwIlbcQ3
— Kirk Borne (@KirkDBorne) 20 décembre 2024
🚀🌟
ChatLLM Teams offers all state-of-the-art #LLMs in one place, and ChatLLM is cheaper than paying for each LLM separately; it’s only $10/month pic.twitter.com/e4utHlCqqpYou can use natural language to generate your #SQL commands using @AbacusAI ChatLLM Teams! Get FREE TRIAL at https://
chatllm.abacus.ai/?token=kirk ChatLLM Teams offers all state-of-the-art #LLMs in one place, and ChatLLM is cheaper than paying for each LLM separately; it’s only $10/month -
Gemini 2.0 Flash Passes Strawberry Challenge Despite Logic Errors
By
–
I just tried the Strawberry challenge on Gemini 2.0 flash and it passed. That said, I'm sure it still makes logic mistakes, including ones that most humans would not make.
-

AI Platforms Analyzing Brand Perception Through LLM Conversations
By
–
With AI platforms like @ChatGPTapp & @perplexity_ai gaining traction, tech firms are trying to understand how LLMs perceive their brands, what outputs mention and what user chat about. The latest example is the AI startup Profound's tool for estimate conversation volume.
