On the @RekaAILabs side, I am thankful that @DaniYogatama gave me the opportunity to experience life as a co-founder when he convinced me to join him 1.5 years ago. I also really enjoyed working with the cracked Reka team (many of them whom are still there), such as @artetxem
@yitayml
-
Incentives Cannot Drive AI Mathematical Reasoning Capabilities
By
–
Great talk by @hwchung27. I really like the interesting analogies here (he's great at that!).
— Yi Tay (@YiTayML) 20 septembre 2024
My favourite one is "no amount of bananas will incentivize monkeys to do mathematical reasoning" π€£ https://t.co/M0iJssLaddGreat talk by @hwchung27
. I really like the interesting analogies here (he's great at that!). My favourite one is "no amount of bananas will incentivize monkeys to do mathematical reasoning" -

GenAI Summit Vietnam Highlights Major AI Talks
By
–
Had a great weekend being part of the GenAI summit in Vietnam. Really enjoyed the impressive talks by @jeffdean
, @quocleix @lmthang and @Diyi_Yang
, along with the nice panel discussion that we had. Meanwhile, one major unexpected highlight of this trip for me personally -
PrefixLM Architecture and Objective: Non-Causal Training Explained
By
–
My take: There's a prefixlm architecture and a prefixlm objective. – prefixlm arch is just a non casual decoder. – prefixlm objective is a standard causal lm training but with a non casual mask before a random split point (to form inputs/targets) during pretaining. The reason
-
Flan-T5 Baseline Remains Competitive for Supervised Tasks
By
–
At small scale, at compute match, with enough supervised data points, a flan-t5 baseline is very tough to beat. Of course t5 has its problems (outdated etc) but a modern one would be very OP.
-
AI Researcher and Engineer Archetypes in Career Development
By
–
Working idea but I've noticed a bunch of archetypes of AI researchers & engineers in my career. Here are some of them: 1. Carry: Hero-level person capable of making unprecedented (alone or in a small group). Either in terms of modeling, infra or making impact in general. Very
-
Model Architectures Discussion on Latent Space Podcast
By
–
I also went on a podcast recently @latentspacepod with @swyx and talked about a bunch of stuff related to model architectures. Check it out here:
-
Model Architecture Deep Dive: Encoders and Prefix Language Models
By
–
Blogpost link: https://
yitay.net/blog/model-arc
hitecture-blogpost-encoders-prefixlm-denoising
β¦ Im gonna write a part 2 and beyond on other topics. I'm interested to know what people find interesting or are dying to know more about. -

Model Architectures in the LLM Era: Transformers and Beyond
By
–
Decided to start a new blog series about model architectures in the era of LLMs. Here's part 1 on broader architectures like Transformer Encoders/Encoder-Decoders, PrefixLM and denoising objectives. A frequently asked question: "The people who worked on language and NLP
-
Evaluating AI Models by Their Real-World Use Cases
By
–
i think the best way to eval these models are how they are going to be used in the end.