The def I'm familiar with is: AGI = "an autonomous system that surpasses human in most economically valuable tasks." For "economically valuable tasks" I like to look at the U.S. Bureau of Labor Statistics index of occupations: https://
bls.gov/ooh/a-z-index.
htm
… You'd have to imagine
@karpathy
-
Defining AGI: Surpassing Humans in Economically Valuable Tasks
By
–
-
Language Model Advantages and Training Data Optimization Strategies
By
–
It's a very amusing thought for sure! Big advantage to languages that have more training data, in both programming (Python / C), and in spoken language (English) 🙂
I'm cautiously optimistic though. Either by dumping docs into contexts, or via synthetic data means. -
A Decade Gap Between Perfect Demo and Commercial Autonomous Driving
By
–
Actually I am very sympathetic to this. Back to my example of self-driving, my first demo drive in an early Waymo was 2014, and it was already great. It took me around for a 20min drive. From that ~perfect demo it was one decade before I could pay for a drive in a Waymo.
-
Automating Software Engineering: Levels of Autonomy and Abstraction
By
–
# automating software engineering
— Andrej Karpathy (@karpathy) 12 mars 2024
In my mind, automating software engineering will look similar to automating driving. E.g. in self-driving the progression of increasing autonomy and higher abstraction looks something like:
1. first the human performs all driving actions… https://t.co/u2EXfxcV6e# automating software engineering In my mind, automating software engineering will look similar to automating driving. E.g. in self-driving the progression of increasing autonomy and higher abstraction looks something like: 1. first the human performs all driving actions
-

Silent Bugs in Fine-tuning: Attention to Detail Matters
By
–
Beautiful work / attention to detail trying to get Gemma to finetune correctly. There are so many foot guns here to be super careful with. All of these issues don't throw any errors, they silently make your network worse. A great example of what I wrote about in my "A Recipe for
-

Training LLMs at Scale: Hidden Infrastructure Challenges and Hardware Health
By
–
Nice read on the rarely-discussed-in-the-open difficulties of training LLMs. Mature companies have dedicated teams maintaining the clusters. At scale, clusters leave the realm of engineering and become a lot more biological, hence e.g. teams dedicated to "hardware health". It
-
Discrepancy between cited and reported accuracy metrics
By
–
Ty for rerunning! Curious btw they cite 84.9 in the release, why is it 82.9 here under “Original”? Maybe you know more about the subtleties here
-
Request for Updated Comparison of Latest AI Models Across Organizations
By
–
Would be interesting to see this updated with the latest models from all orgs
-

Claude 3 Tokenization Challenge: Stylistic but Contains Subtle Hallucinations
By
–
Claude 3 takes on the Tokenization book chapter challenge 🙂 context: https://t.co/yRaeTbkblY
— Andrej Karpathy (@karpathy) 4 mars 2024
Definitely looks quite nice, stylistically!
If you look closer there are a number of subtle issues / hallucinations. One example there is a claim that "hello world" tokenizes into 3… https://t.co/mMjK27Uk2yClaude 3 takes on the Tokenization book chapter challenge 🙂 context: https://
x.com/karpathy/statu
s/1760740503614836917
… Definitely looks quite nice, stylistically! If you look closer there are a number of subtle issues / hallucinations. One example there is a claim that "hello world" tokenizes into 3 -
Understanding Consistent Behavior Across Independent Language Models
By
–
I’d love to understand this better too… I thought it was just a quirk of the specifics of labeling instructions, but then multiple (what I think should be mostly independent) language models seem to all do this.