"because after the event, everyone is a sage. You are doing the same." No, friend, it doesn't work like that. Here, at every point along the way for several years, we have had people denying that the models would improve further, seeing invisible walls. If it had depended on them, they would not have…
@dotcsv
-
The nonexistence of the wall in AI model improvement
By
–
The wall as a concept that we can predict and place at a specific point in time where model capabilities will not improve does not exist, and none of the attempts to draw that boundary in the last 3 years have been successful. It is another thing to say that at some
-
Fable/Mythos demonstrates the absence of a wall in the evolution of LLMs
By
–
The most important thing about the release of Fable/Mythos is the demonstration, once again, that there is no wall. The paradigm of LLMs and test-time-compute, far from slowing down, continues to climb the hill by delivering increasingly intelligent and capable models,
-

A model more capable than a year ago, priced below Opus 4.1
By
–
If you really look at it, this model, much more capable than what we had a year ago, is released with a price below Opus 4.1.
-

Model now limited, previously considered dangerous
By
–
"pEr0 Sí hAcE 2 MeSeS deCIAn eRa P€ligRos0…" Leaving aside the fact that 2 months ago they didn't have enough computing power to release it, IN ADDITION, the model they released today is limited, precisely solving these security issues. You have read
-
The AI bubble is postponed, we continue to inform.
By
–
The AI bubble is postponed. We will continue to inform.
-

Anthropic would limit capabilities to maintain competitive advantage
By
–
Pretty crazy this that is being shared where Anthropic would be limiting the model's capabilities when used to improve and create better LLMs. They sell it as a security measure but it is clear that they do it to maintain their competitive advantage.
-

Mythos break expected progress line in System Card
By
–
Returning to the System Card, here is one of my favorite graphs, and it shows how the Mythos category models break the expected line of progress. As always with AI, we reach higher levels ahead of time.
-

Cost frontiers: more expensive and superior model, Fable surpasses Opus 4.8
By
–
The cost frontiers, as expected, show an upward and rightward shift, meaning we have a more expensive and notably superior model (the low of Fable remains well above the x-high of Opus 4.8)
-

Benchmark tables: the model dominates in everything evaluated
By
–

Here are the complete benchmark tables. The model basically dominates in everything evaluated in these tables (Humanity Last Exam, CriPT, ArxivMath, HealthBench, etc.)