Having played with it a bunch, Horizon-alpha did a pretty solid version of Missile Command with Relativistic Effects with a few rounds of feedback, passed the Lem Test the first time (without reasoning), and drew a passable TikZ unicorn (if you know, you know). Very quick model.
LLMS
-
Free Step-by-Step AI Agents and RAG Systems Tutorials
By
–
100+ free step-by-step tutorials with code covering: AI Agents RAG Systems Voice AI Agents MCP AI Agents Multi-agent Teams Autonomous Game Playing Agents P.S: Don't forget to subscribe for FREE to access future tutorials.
-
Perplexity Finance reliability compared to 3B language model
By
–
perplexity finance’s numbers have the reliability of a 3b llm btw still unserious (for now)
-
Claude Artifacts Now Enable AI-Powered Apps with API Access
By
–
The second is the ability for artifacts to build "AI-powered apps" to which users can upload PDFs and images
— Simon Willison (@simonw) 31 juillet 2025
What that means: Artifacts can call the real Claude API now via a monkey-patched fetch() that gets redirected through a proxy adding an API keyhttps://t.co/V5VJeR6SaQThe second is the ability for artifacts to build "AI-powered apps" to which users can upload PDFs and images What that means: Artifacts can call the real Claude API now via a monkey-patched fetch() that gets redirected through a proxy adding an API key
-

Claude Model Cannibalization: Token Usage Trends Sonnet Releases
By
–
How much are Claude models cannibalising themselves? After each new release, usage of the previous model declines, yet the overall consumption increase, e.g., weekly @openrouter usage rose from 211b tokens for Sonnet 3.7 to 287b tokens for Sonnet 4 (a one-third increase).
-

AI Model Performance Assessment on Coding Tasks
By
–
This is quite good though But tested on some coding tasks, didn't seem that much better
-

Mixture-of-Recursions: Adaptive Computation for Language Models
By
–
What if language models could learn to "think harder" only when they need to—allocating deep computation to challenging tokens while breezing through simple ones? We're excited to have Reza tomorrow in our AI4Science community giving a talk on Mixture-of-Recursions!
-
Weaver Combines Weak Verifiers to Boost LLM Output Quality
By
–
New from Snorkel AI: Weaver combines weak verifiers to boost LLM output quality: no fine-tuning needed. +14.5% accuracy 98.2% of gains, 99.97% less compute (via distillation) Outperforms naive voting, rivals GPT-4o Read the full paper
-
AI Model Excels Across Languages and Receipt Recognition
By
–
Does super well on that too, any language too, including receipts.
-
KL Divergence Measurement on Reinforcement Learning Outputs
By
–
no, because it's KL measured on the RL outputs
