lack of education perhaps (which is what my tweet is supposed to mitigate).
@yitayml
-
2023 Challenges for AI Researchers: A Critical Analysis
By
–
Hot take: 2023 is not a good year for being an AI researcher.
-
UL2 Training Objective: Notable Models and Research Papers
By
–
It’s been slightly more than a year since the UL2 paper (
https://
arxiv.org/abs/2205.05131) was released. Here’s a summary thread of some notable models/research papers that use the UL2 objective for training (aside from the original UL2/Flan-UL2 of course). thread below #1 – -
ChatGPT Playing Pokemon Showdown and Video Game Emulators
By
–
i would love to see chatgpt play pokemon showdown or some emulator
-
LLM Evaluation and Benchmarking: Critical Missing Piece
By
–
Eval and good benchmarking for LLMs is one very big piece missing right now.
-
Tracking Emergent Abilities and Capabilities in AI Systems
By
–
emergent abilities or capabilities track?
-
Research Organization Evolution: From Applications to New Paradigms
By
–
Just a few years ago, research is mostly sorted by "applications". When folks asked what research you're working on, you're expected to say something like "oh I work in question answering" or "sentiment analysis" or something . In fact, all the conference tracks are sorted as
-
Hype-Driven Claims: Short Fame, Long Cleanup for Science
By
–
One minute of fame for hype grabbers. One year to clean up the mess for the community. Negative impact in science exists unfortunately
-
Language Model Evaluation Tasks and Benchmarking Methodologies Comparison
By
–
Oh yeah. I know the {0,1,N} shot tasks in LM harness and in the palm/gpt-3 evals are very similar modulo some prompting diffs. I don't exactly mean to say palm-evals are better than that. It was just referring to the academic tasks in general (not specifically LM harness). Im
