Producing many coherent answers to an open-ended question without repeating is appealing IMO because it both uses world knowledge and it demands a model pay attention to everything in context. And there’s no ceiling.
LLMS
-

o3-mini achieves #1 and #2 on Aidenbench, topping big model smell
By
–
o3-mini takes #1 and #2 on Aidenbench A mini now tops the benchmark of “big model smell.”
-
o3-mini-high Model Quality and AI Naming Strategy Challenges
By
–
o3-mini-high is really good and it's cool to hear some people say it's their favorite model ever but it reminds me of when we used to have something internally called "the little big run" a top 2025 goal is to fix our naming problem
-
o3-mini Successfully Solves Problem on First Attempt
By
–
So cool, o3-mini nailed this one on the first go! https://t.co/VjGlSCWXE1
— Romain Huet (@romainhuet) 1 février 2025So cool, o3-mini nailed this one on the first go!
-

FinRobot: AI Agent for Equity Research and Valuation
By
–
FinRobot: AI Agent for Equity Research and Valuation with Large Language Models https://
bit.ly/3AY2daK
#AI #MachineLearning #DeepLearning #LLMs #DataScience -

Agent Reasoning Interface: o1/o3, Claude 3, ChatGPT Canvas
By
–
The Agent Reasoning Interface: o1/o3, Claude 3, ChatGPT Canvas, Tasks, and Operator ft @karinanguyen who is teasing her keynote at @aiDotEngineer Timestamps •00:11 Introducing Karina Nguyen
•02:21 Karina's Journey to OpenAI
•04:45 Early Prototypes and Projects
•05:25 -
o3 mini Function Calling Solves Complex Math Science Problems
By
–
What we love about o3 mini is the ability to use function calling.
— FlowiseAI (@FlowiseAI) 1 février 2025
With WolframAlpha and Code Interpreter, o3 mini is able to solve science and math problems which gpt4 had failed to do so.
It can tackle questions that involve combining several calculations and formulas 👇 pic.twitter.com/UbxjyCW5JVWhat we love about o3 mini is the ability to use function calling. With WolframAlpha and Code Interpreter, o3 mini is able to solve science and math problems which gpt4 had failed to do so. It can tackle questions that involve combining several calculations and formulas
-

ChatGPT Auto Mode: Model Selection Based on User Questions
By
–
There should be an “Auto” mode where ChatGPT picks the model based on the question it’s asked
-
LLM usage limited like calculator, capable of extreme potential
By
–
Most people use LLMs with prompts that are the equivalent of putting “2+2” into a TI-84 calculator.
— AI Breakfast (@AiBreakfast) 1 février 2025
They are capable of so much more and differentiate themselves in the extremes. https://t.co/3NrKXPlDWeMost people use LLMs with prompts that are the equivalent of putting “2+2” into a TI-84 calculator. They are capable of so much more and differentiate themselves in the extremes.
-

o3-mini beats R1 and o1 on Humanity’s Last Exam
By
–
o3-mini beats R1 and o1 on Humanity’s Last Exam (big asterisk for the latter, though; problems selected to be hard for existing LLMs)