Think most people work on LLM RL to some extent these days.
@yitayml
-
Sota Singaporean AI Researcher List Update with New Names
By
–
Sota Singaporean AI researcher list I made some list two years ago and thought it's time to revisit it with some new names (in no order and no elaborations/cot this time) 1) @isaacongjw from Anthropic
2) @alyssamloo from Anthropic
3) @jennyzhangzt from UBC
4) @tanshawn from -
Life Advice Explained Through Machine Learning Terminology
By
–
I found out that I am pretty good at giving life/relationship advice to others. The only problem is that I will explain it to you in ML terminology. https://
x.com/_jasonwei/stat
/_jasonwei/status/1940126761489928468
… -

Google Releases Gemini 2.5 Flash-Lite Model
By
–
Gemini! The new 2.5 flash-lite has the same crazy research we landed in the 2.5 flash launched at google i/o. There's also a Gemini 2.5 tech report now.
-
Hundreds of AI Model Checkpoints from Scaling Research
By
–
I now realized I probably have dozens of models there (maybe a hundred?) due to a dump of many checkpoints from one scaling paper, plus ul2 and flan-t5/ule series. And also umt5.
-
Generative Retrieval: Making IR Cool Again at Google
By
–
Oh we're doing more fun BTS here? I accidentally invented generative retrieval as a side project because my manager at Google at that time was very into IR/retrieval but IR was kind of a sunset/boring field so I kind of wanted to make it cool again After DSI was born,
-
Dataset Collection Challenges and Academic Impact in AI Research
By
–
We had fun though I guess! But collecting all those datasets was indeed painful and I feel a little bad for all the contributors that we could not "convert" their effort well enough into a few K citations paper. My bad for the 200 citations consolation prize.
-
Research Motivation Impact on Project Outcomes and Lessons
By
–
Sharing a pretty interesting story about how two research projects can end up with a very different outcome despite almost similar type of “work” being done, just because of research motivation and taste. Here, I was on the side that got rekt so I thought I would share my lessons
-
Model trainers as RL agents: unprincipled YOLO runs reconsidered
By
–
i don't think yolo runs are actually really "yolo". model trainers (humans) are agents that are "RL-ed" by the environment of running many experiments and getting a lot feedback through their careers. it's becoming one with the model. it's not unprincipled, we just don't
