Benchmarks:
– MMLU (massively multitask language understanding): https://
arxiv.org/abs/2009.03300
– BBH (Big-Bench Hard): https://
arxiv.org/abs/2210.09261
– TyDiQA (typographically diverse QA): https://
arxiv.org/abs/2003.05002
– MGSM (multilingual grade school math): https://
arxiv.org/abs/2210.03057
@_jasonwei
-
Key AI Benchmarks for Language Model Evaluation
By
–
-

Text-davinci-003 Instruction Following vs Academic Benchmark Performance
By
–
@OpenAI
's text-davinci-003 follows instructions better. Is it also better on academic benchmarks? Summary:
– text-davinci-3 beats text-davinci-2, but is not as good as code-davinci-2
– it is behind @GoogleAI
's PaLM and Flan-U-PaLM Full results: https://
arxiv.org/abs/2210.11416 App D -
Code-Davinci-2 vs Text-Davinci-3: Instruction Tuning and PPO Performance
By
–
– code-davinci-2 > text-davinci-3 means that their instruction finetuning overall hurts performance on academic benchmarks
– text-davinci-3 > text-davinci-2 means that PPO improves performance -
Text-davinci-003 results now available in paper appendix
By
–
https://
arxiv.org/abs/2210.11416 Appendix has text-davinci-003 results now -
NeurIPS 2022: Meet Jason Wei at Google booth and CoT poster
By
–
i'm at neurips, would love to chat!
– google booth, 10:30-11am, 1-2:30pm tues
– poster on CoT prompting, 11am-1pm wed
– let's meet up, dm me! -
Old NLP Experts Underestimate Language Models’ Capabilities
By
–
Some old-time NLPers can be closed minded about the "range of tasks" that LMs can do. They are still thinking about LMs through BERT, etc. I rarely see this issue with new-joiners to NLP.
-
Correction: Chain of Thought paper accepted at NeurIPS
By
–
Oops i had a typo, CoT is in NeurIPS not ICLR
-
Research Papers and PhD Blog Post Recommendations
By
–
These are the papers: https://
arxiv.org/abs/2201.11903 https://
arxiv.org/abs/1901.11196 https://
arxiv.org/abs/2210.11416 https://
openreview.net/forum?id=gEZrG
CozdqR
… Also, i saw this from Andrej Karpathy's blog: http://
karpathy.github.io/2016/09/07/phd/ -

Dessert-themed models preferred over DeepMind’s rodent research
By
–
enjoying the dessert themed models. i like deepmind but desserts are cooler than rodents, sorry
-
Large Language Models Should Master Multi-Digit Addition
By
–
Multi-digit addition should definitely be within reach for large language models at this point! https://t.co/snqwqxnVQ7
— Jason Wei (@_jasonwei) 18 novembre 2022Multi-digit addition should definitely be within reach for large language models at this point!