MedGemma is a really interesting model – very small, multimodal, open, and does quite well in out-of-distribution medical tasks compared to much larger models. Would love to see more work thinking about how to improve & deploy this sort of LLM to support medical professionals
LLMS
-
Long-form coding tasks with AI models
By
–
Wow exciting stuff! 😀 Have you tried it with long-form coding tasks?
-

Upstage Launches Solar-Pro2 Frontier Model with Advanced Reasoning
By
–
Finally, @upstageai is proud to announce the launch of its #Frontier model #SolarPro2. We sincerely thank everyone who used our solar-pro2-preview and provided valuable feedback. Leveraging your input, we have incorporated newly developed learning methodologies, enhanced data curation techniques for improved data quality, general reasoning data synthesis methods, and verifiable reward technologies to create this advanced model. The Solar-Pro2 model demonstrates exceptional performance on high-difficulty reasoning benchmarks, including MMLU-Pro, Math500, AIME, and SWE-Bench. Moreover, it excels in Arena-Hard-Auto, a rigorous evaluation framework adapted from LMSYS’s LLMArena, which focuses exclusively on extremely challenging problems. Arena-Hard-Auto is not a simple multiple-choice test—it evaluates answers against those generated by top-tier models, akin to lengthy essay or subjective exams. This benchmark is so demanding that non-Frontier models often score only 10–20%, making it nearly impossible to publish results. Solar-Pro2 (reasoning) achieves a 46% score on Arena-Hard-Auto and 71% on Ko-Arena-Hard-v0.1, surpassing even the highest-quality Korean answers. Solar-Pro2 is now available for immediate use at chat.upstage.ai/ and via API (model name: solapr-pro2). To encourage broader adoption, we offer a TRY-SOLAR-PRO2 $50 credit code (valid until August 31).
-
100M Parameter Networks Matching O3 Pro in 100 Years?
By
–
in 100 years, will we have 100M parameter neural networks that perform at the level of today's o3 pro?
— dr. jack morris (@jxmnop) 9 juillet 2025
and if so, how? https://t.co/i80ZeIH3Kyin 100 years, will we have 100M parameter neural networks that perform at the level of today's o3 pro? and if so, how?
-

Compressing Knowledge: Can Small LLMs Match o3 Pro Capability?
By
–
more on this idea in our pod with @jxmnop
— swyx 🐣 (@swyx) 9 juillet 2025
but his favorite part is highlighted here: the "cognitive core" question of how smol we can compress knowledge and LLMs so that they have o3 pro level capability with only x00m parameters?https://t.co/4Skn4AKlgu pic.twitter.com/dFhlToWKC9more on this idea in our pod with @jxmnop but his favorite part is highlighted here: the "cognitive core" question of how smol we can compress knowledge and LLMs so that they have o3 pro level capability with only x00m parameters?
-

MedGemma: Open Weights Multimodal Model for Medical Data Analysis
By
–
Check out our state-of-the-art open weights MedGemma multimodal model for making sense of longitudinal EHR data as well as medical text and medical imaging data in various modalities (radiology, dermatology, pathology, ophthalmology, etc.) See the blog post linked below!
-
o3 Research AI Model Parallel Verification Cross Check
By
–
o3 is just so sublime on research, I can't get enough of using it. If I need something more serious, I'd throw in parallel requests to o3 in multiple windows and cross check. It is a beast
-

Do Language Models Think Differently in Other Languages?
By
–
Do language models 'think' differently in other languages? The answer seems to be 'no'. When asked 'What's your favourite number?' in 30 different languages, models answered '7' 90% of the time and '42' 9% of the time. This mirrors the results we had with the question 'What's
-

Claude 4.1 and 4.2 in Development, Llama 5 Market Prediction
By
–
apparently 4.1 and 4.2 are in the works still! yay! https://
x.com/ldjconfirmed/s
tatus/1943030722329219484
… but my polymarket is on llama 5 vs rebrand
