Here I'm assuming we'll have AI so good it will be able to do what a human driver can do, with no new downsides. In practice, there will be various new unit costs, including the labor cost of remote monitoring/operation.
@fchollet
-
Labor costs represent only 20-40% of ride prices, not 90%
By
–
You seem to assume that labor cost is 90% of the price of a ride. In reality it's 20-40%. Ride price won't drop much (there are also new costs centers that arise). In fact Waymo is currently more expensive than Uber in SF.
-
Human Intelligence is Compositional, Not Data-Dependent
By
–
Needless to say, this is not how human intelligence works. Human intelligence is compositional, which means you can understand the cross product of two spaces without being explicitly exposed to a dense sampling of data pairs from those spaces. When reading a book, most people
-
Why Frontier VLMs Lag Behind LLMs and Vision Models
By
–
Frontier LLM have superhuman text-based world knowledge. Frontier image / video models have superhuman vision-based world knowledge (e.g. Genie). But current frontier VLMs are still absolutely clown shoes. Why? Relative scarcity of image:text pairs (while there is plenty of text
-

JAX Performance Meets Keras 3 Developer Velocity
By
–
JAX = performance & scalability Keras 3 = high velocity development, compact code, best practices by default Both at the same time = pretty killer
-
AGI Timeline Uncertain, Won’t Come From Current System Scaling
By
–
AGI might happen soon-ish, but won't be coming from scaling up current systems, which makes it tricky to time — definitely not a matter of extrapolating from a chart
-
ARC-AGI v1 eval sets difficulty calibration and v2 improvements
By
–
Note that the v1 semi-private eval set is about the same level of difficulty as the v1 public eval set. The fully private eval set (used for the 2020 and 2024 Kaggle competitions) is estimated to be more difficult. ARC-AGI-2 does not have this issue — all sets were calibrated
-

Qwen3-235b Instruct Verified: 11% ARC-AGI-1 Performance, Most Cost-Effective
By
–
Official verification of Qwen3-235b Instruct: it gets 11% on ARC-AGI-1 and 1.3% on ARC-AGI-2 (semi-private sets). These numbers are in line with other SotA base models. Qwen3 stands out by being the cheapest base model we tested to score above 10% on ARC-AGI-1.
-
Resist the tendency to anthropomorphize non-human things
By
–
Resist the tendency to anthropomorphize that which is not human
-
Qwen 3 ARC-AGI score cannot be reproduced independently
By
–
Please note, we're not able to reproduce the 41.8% ARC-AGI-1 score claimed by the latest Qwen 3 release — neither on the public eval set nor on the semi-private set. The numbers we're seeing are in line with other recent base models. In general, only rely on scores verified by