Nice paper by Tu Vu on factuality in LLMs: http://
arxiv.org/abs/2310.03214, enjoyed contributing in a minor role to it while I was at Google. The main takeaway for me is that most factuality benchmarks for LLMs don't really take into account the fact that many types of knowledge
@_jasonwei
-

Factuality in LLMs: Benchmarks and Knowledge Assessment
By
–
-

Beating the Scaling Curve: Research, Engineering, and Leadership
By
–
Pretty rare to see a single talk cover three main areas of AI (research, engineering, and thought leadership), but my colleague @hwchung27 has experience and original takes on all of them. Some notes on my takeaways: (1) To “beat the scaling curve”, always think about how
-
Reaching 10k Citations: Reflections on Research Career
By
–
I reached 10k citations recently, a goal of mine for many years. It’s a nice moment to reflect back, and I mostly feel bittersweet: (1) When I joined Google Brain back in 2020, I thought I'd stay for 10+ years, doing open-ended research and publishing papers. But the field has
-
Quick Release Support for Audio Video Text and Multilingual
By
–
Cool to see this come out so quickly, and already supports audio, video, and text as well as multilingual. Looking forward to playing with it once it is released. Congrats @YiTayML
! -
Great AI Researchers Manually Inspect Data for Valuable Intuitions
By
–
One pattern I noticed is that great AI researchers are willing to manually inspect lots of data. And more than that, they build infrastructure that allows them to manually inspect data quickly. Though not glamorous, manually examining data gives valuable intuitions about the
-
Transparency on Social Media: Anonymous vs Named Accounts
By
–
Some have noticed that I am quite transparent and personal on Twitter/X. I had an interesting discussion with some of the prolific anonymous accounts (
@alth0u and @tszzl
) about anonymous versus named accounts. My recent thoughts: (1) The most-cited benefit of anonymous accounts -
Language Model Evaluation Methods: From Casual Testing to Rigorous Benchmarks
By
–
Just a joke, don’t take this meme too seriously and pls do rigorous evals 🙂 Explanation:
– Left: The simplest way to evaluate a language model is to play with it for 15 minutes. This is not scientific at all.
– Middle: The more systematic way is to create a diverse set of -

Language Model Evaluation: Challenges and Solutions After 5 Months
By
–
Large language models are notoriously hard to evaluate because (1) they are highly multi-task, (2) they generate long completions, and (3) grading is subjective. After spending ~5 months rigorously working on how to do language model evals, this is my verdict:
-
Amusing Nuggets from My Google Brain AI Residency Work Log
By
–
I recently dug up my work log from when I was an AI Resident at Google Brain and found some amusing nuggets: – I used to block out specific periods of time to code while wearing a full three-piece suit. (Yes, wearing the suit did improve my productivity.)
– Every Saturday night, -
Considering Long-term Utility of Fine-tuned Models
By
–
Agree it's good to use what's best right now, but my personal opinion (and glad to hear arguments for why you think predictions are unneeded) is that it's important to think about how long something you finetuned will be useful. Maybe it's easy to finetune something in a few