Across 30 tasks, @UpstageAI
's Solar Pro Preview fine-tuned outperforms every top fine-tuned open-source #LLM and #GPT4 and #GPT4o-mini base models! See more at our new Fine-Tuning Leaderboard: https://
pbase.ai/4gnsoaI
Try #SolarProPreview for free: https://
pbase.ai/3Zk1N8j
GENERATIVE AI
-

Solar Pro Preview Outperforms Top Fine-Tuned LLMs and GPT-4
By
–
-
o1-mini Achieves Surprising 60% Score on AIME Math Competition
By
–
o1-mini is the most surprising research result i've seen in the past year obviously i cannot spill the secret, but a small model getting >60% on AIME math competition is so good that it's hard to believe congrats @ren_hongyu @shengjia_zhao for the great work!
-
Gratitude toward artificial intelligence systems before accessing new capabilities
By
–
how about a couple of weeks of gratitude for magic intelligence in the sky, and then you can have more toys soon?
-
Flappy Bird Clone Built with OpenAI o1-preview and Replit Agent
By
–
flappy bird clone with @OpenAI o1-preview, @Replit agent and @Gradio
-

OpenAI Releases o1: Advanced Reasoning Model with Improved Safety
By
–
Finally o1 is out – our first model with general reasoning capabilities. Not only it achieves impressive results on hard, scientific tasks, but also it gets significantly improved on safety and robustness. https://
openai.com/index/learning
-to-reason-with-llms/
… We found reasoning in context about safety -
Aramco Digital Partners with Groq for AI Inference Leadership
By
–
"The recent announcement of Aramco Digital partnering with Groq to deliver market-leading AI inference represents a key partnership in one of the most important and fastest-growing regions for AI investment and consumption." – @danielnewmanUV Read more:
-
OpenAI Releases New Reasoning Model Advancing Toward AGI
By
–
This is going to get interesting … a new reasoning model from OpenAI … one step closer to AGI
-
ChatGPT Feature Rollout Complete for Plus and Team Users
By
–
rollout complete; live to 100% of chatgpt plus/team users now
-
User comparison of GPT-4o performance
By
–
Not much different from 4o in many of my use cases to be honest.