sam altman wants to build a model that’s rated 3200 on codeforces but has no idea who Sam Altman is and i think that’s beautiful
LLMS
-
GPT-OSS Model Performance: Coding Excellence Mixed with Factual Hallucinations
By
–
i’ve spent the last couple hours talking to gpt-oss and can safely say it’s unlike any model i’ve tested one second it’s coding for me at a professional level, the next it’s making up basic facts and clinging to them no matter what i say something very strange is going on
-

GPT-OSS Models Reach #1 and #2 on Hugging Face
By
–
Both gpt-oss models are trending #1 and #2 among 2M models on @huggingface
! Thanks to the open-source AI community for your support since launch. We’re following discussions and will pop in when we can—feel free to ask questions, share ideas, and show what you’re building! -
GPT-5 Achieves 90% on Public Benchmarks, Promising Results
By
–
Once again about the presumed GPT-5 benchmarks: testing was done with the public simpleBenchDatasets. These indicate that around 90% is achieved. That is excellent – and in all probability GPT-5 will be an outstanding model. I don't think there's any doubt about that now.
-
Clarification on GPT-5 Testing and Public Dataset Information
By
–
Edit: just to make it clear: I don’t know if they used the public set to test on simple bench and if it was really GPT-5. don’t wanna spread misinformation. That being said it doesn’t mean it’s untrue. I just can’t confirm anything.
-
RULER: General Purpose Reward Function by OpenPipe AI
By
–
I think you're describing @OpenPipeAI
. Check out their work on RULER (
https://
openpipe.ai/blog/ruler?ref
resh=1754513766765
…), it's essentially a general purpose reward function. Might want to chat with them! -
GLM-4.5 Air Outperforms OpenAI 20B on Space Invader Prompt
By
–
I tried the same space invader prompt against the OpenAI 20B model and thought the GLM-4.5 Air result was better – though that model is 3x the memory size of OpenAI's
-
GitHub Copilot Launches Asynchronous Coding Agent Feature
By
–
I also spotted it in the press release for GitHub's own implementation of this pattern, the somewhat vaguely named "Coding Agent for GitHub Copilot", which they introduce with the line "GitHub Copilot now includes an asynchronous coding agent": https://
github.com/newsroom/press
-releases/coding-agent-for-github-copilot
… -

GPT-5 in Copilot Won’t Drive User Switching
By
–
Despite GPT-5 in Copilot, nobody will probably switch anyway.
-

OpenAI Major Announcement Generates Excitement Among Developers
By
–
Just got this – it's like Christmas, Birthday & birth of your first child happening same week @OpenAIDevs @edwinarbus