I am not garymarcusing here. LLMs are proof that it is possible to distill intelligent reasoning behavior by doing statistics over human language patterns
RESEARCH
-

Learning world models from dropping a cup of water
By
–
When @olivercameron was the first to teach us about world models (he just collected $300 million investment last week) in my head I was thinking: "If I drop a cup on the ground, with some water in it, and film that with a high speed camera, the world model would learn a lot about
-
Frontier labs’ billion-dollar brute-force data annotation with diverse annotators
By
–
wish there was more public info on what's happening behind the scenes. frontier labs are spending BILLIONS paying {poets, musicians, accountants, consultants, …} to annotate massive amounts of data:
• essays
• slides
• spreadsheets
it's a brute-force bet. but seems to be -
Anthropic Mythos new high-performance version, only the beginning
By
–
A new, more powerful version of Anthropic's Mythos has come out of training. In itself, this is nothing extraordinary. What else could one expect? That Mythos is already the end? Of course not. This is only the beginning. What is exciting here is the
-
Anthropic’s Fable and Mythos risks highlight Europe’s AI giants gap
By
–
The examples around Fable from Anthropic, risks from Mythos, lack of AI giants across Europe including UK leads to gaps in defence and economics
-

Equivalent SOTA model, open access, not free, no third-party dependency
By
–

In fact, you have a model equivalent to SOTA models on many benchmarks, freely accessible, obviously not free because you have to account for the cost of hardware, but thus possibly without depending on a third-party actor. What is good when a new model comes out is to look at
-
Comparing models is not limited to pure performance
By
–
Comparing models is not just a matter of pure performance.
-
Issues with Codex and Code: division, lack of exploration and testing
By
–
This is aside from the other key "software brain" problems of Codex and Code: dividing all work into front-end and back-end design, solving for the general case in a repeatable way, not testing or exploring idea spaces, testing for technical correctness but not other aspects…
-

Top AI Papers of the Week: June 14-21 Highlights
By
–
The Top AI Papers of the Week (June 14 – June 21): – PreAct
– SpatialClaw
– Back on Track
– OpenClaw-Skill
– From Trainee to Trainer
– Compositional Skill Routing
– Can LLM Agents Infer World Models? Read on for more: