This new hallucinations eval by GDM friends is in the right direction in many ways: 1. Tackles the scenario of extremely long-form responses, which is a harder but more realistic setting
2. Extracts the number of relevant facts, then browses to verify each individual fact
3.
@_jasonwei
-

New Hallucinations Evaluation: Long-form Responses and Fact Verification
By
–
-
AI Development as the Great Power Competition of Our Era
By
–
Cheesy realization: studying history underscores how special this current moment in AI is. In past eras, the great powers of the world fought religious wars, sailed to unexplored lands, and built the first industrial cities. Now we will race to build artificial intelligence. So
-
Sora: The GPT-2 Moment for Video Generation Technology
By
–
My mental model of Sora is that it is the “GPT-2 moment” for video generation. GPT-2, which came out in 2018, could generate paragraphs of text that are coherent and grammatically correct. GPT-2 wasn’t able to write an entire essay without making mistakes like being inconsistent
-
A Day in the Life of an OpenAI Technical Staff Member
By
–
My typical day as a Member of Technical Staff at OpenAI:
[9:00am] Wake up
[9:30am] Commute to Mission SF via Waymo. Grab avocado toast from Tartine
[9:45 am] Recite OpenAI charter. Pray to optimization Gods. Learn the Bitter Lesson
[10:00am] Meetings (Google Meet). Discuss how to -
Congratulations on AI Vision and Passion Launch
By
–
Congrats! I have been inspired by your vision and passion for AI since we worked together in the late 2010s, which now feels like so long ago. Super happy to see this exciting launch 🙂
-
The Art of YOLO Runs in AI Research at OpenAI
By
–
An incredible skill that I have witnessed, especially at OpenAI, is the ability to make “yolo runs” work. The traditional advice in academic research is, “change one thing at a time.” This approach forces you to understand the effect of each component in your model, and
-
Congratulations on Being First to Open-Source Model Parameters
By
–
Wow, super cool, congrats! Only closed-source company to open-source how man params your model is 🙂
-
Chain-of-Thought Information Density and Compute Scaling
By
–
A key insight from chain-of-thought is around the idea of information density. Language models can only do so much with a single forward pass, and so the amount of compute the language model can use must be scaled proportional to how hard a prompt is to solve. What is
-
Typo in launch script ruins overnight experiments
By
–
One of the great pains in life is waking up and finding out the experiments i launched last night had a typo in the launch script
-
The Joy of Checking Overnight Experiment Results
By
–
One of the great pleasures in life is waking up and immediately going to my computer to check the results of experiments I launched last night