Excited for more open science with LLM frontier research and deployments!
LLMS
-

Perplexity Open-Sources R1 1776 Model
By
–
Perplexity has open-sourced R1 1776 and its weights are now available on Hugging Face. "R1 1776 is a DeepSeek-R1 reasoning model that has been post-trained by Perplexity AI to remove Chinese Communist Party censorship." https://x.com/perplexity_ai/perplexity_ai/status/1891916573713236248
-
Good Stories Take Time: Rachel Metz on LMARENA AI
By
–
Sometimes good stories take a while to finish cooking. I started this one last summer, and the frenzy over DeepSeek provided a great news hook to get back into it. My latest for @BW
, on @lmarena_ai
. Gift link! -
Technological Overhang: More Science Than We Can Apply
By
–
It's a really strange time. We've never had a technological overhang like this. We have more science than we know how to apply. Every day people are still finding so many new amazing things to do with LLMs week after week.
-

New Coding Benchmark Released, But Uses Older Sonnet Version
By
–
Cool new coding benchmark! I always love to see new evals out in the world. Note though that the testing here is on the June version of Sonnet, not the latest version, so technically not "current frontier models."
-

Frontier AI models struggle with majority of tasks
By
–
Current frontier models are unable to solve the majority of tasks.
-
SWE-Lancer: New AI Coding Performance Benchmark Launched
By
–
Today we’re launching SWE-Lancer—a new, more realistic benchmark to evaluate the coding performance of AI models. SWE-Lancer includes over 1,400 freelance software engineering tasks from Upwork, valued at $1 million USD total in real-world payouts.
-
GPT-5: An AI Dropdown Menu to Select the Right Model
By
–
Plot-twist: GPT-5 is an AI-powered dropdown menu that picks the right model for the task
-
OpenAI o3-mini and o3-nano Models Enthusiasm
By
–
the more the merrier: o3-mini and o3-nano would be cool!
-
Request for Recording and Technical Details on Reasoning Model Methodology
By
–
Thanks for the comprehensive summary!! Unfortunately, I wasn't able to tune in to the event but was there a recording or technical blog/paper with info on their reasoning model methodology (re RL, inference-compute scaling strategies etc)?