academic tasks, doesn't reflect real world use. models scoring well on helm doesn't necessary mean they are good (or vice versa).
@yitayml
-
Human raters blindly evaluate AI model outputs
By
–
referring to the fact that the human raters don't know where (which model) the outputs come from.
-
Systematic Blind Evaluation Benchmarks for LLMs Urgently Needed
By
–
friendly reminder to everyone that there isn't yet a good & proper systematic blind eval/benchmark of LLMs yet, especially those on real world data/use-cases. if i were in academia this is something i'll work on immediately.
-
FLAN-UL2 Not a Chatbot Understanding Model Purpose
By
–
IIRC, when it first came out i saw some people trying to compare it with chatgpt. flan-ul2 is not a chatbot, folks who already use it know where it's true value lies. some ppl be lik going to supermarket, buying raw chicken breast, eating it and complaining it's not KFC .
-
Proprietary Model Dependency Risk in AI Development
By
–
Yea, one group's cash-in will hinder progress substantially. Everyone will use api-distilled dataset to bootstrap their models. 5 years down the road…no one figured how to do it without the proprietary model in the first place. So dangerous.
-
Emergent Abilities in Large Language Models Discussion
By
–
Had fun talking to @QuantaMagazine about emergent abilities in LLMs. https://t.co/IawuauWrmk
— Yi Tay (@YiTayML) 17 mars 2023Had fun talking to @QuantaMagazine about emergent abilities in LLMs.
-
Worker Exhaustion Linked to Intensive Labor Demands
By
–
Maybe people are exhausted because they are working hard
-
PhD Students and Academia Have Abundant Research Opportunities
By
–
There are still tons of stuff that PhD students and academia can work on!
-
Flan Fine-tuning Hyperparameters Same as Flan-T5
By
–
for flan finetuning? the hparams are same as flan-t5 https://
arxiv.org/abs/2210.11416.. -
Google PaLM API Preview Released for Developers
By
–
Take a sneak peak at what the PaLM API looks like https://
t.co/fru1rungcx