Now this is the type of excellent work that the community needs more of! Everyone cranking out minute-of-fame "model distillation" papers should take a look at this fine exemplar of good science below: great work @EdwardSun0909
@yitayml
-
Is the new LLM hype cycle just a mirage?
By
–
Maybe all the hype of all these new LLM models ……is a mirage?
-

RNN Cells: A PhD Research Retrospective on Recurrent Architecture
By
–
Ah, RNN cells. I seldom share the work I did during my PhD (pre-Google) but I thought maybe people would find this work fun and amusing. https://
arxiv.org/abs/1811.09786 Throwback to the fond memories of drawing recurrent cells in papers. If you look into that paper you will see that -
Paper acceptances don’t matter for spot bonuses
By
–
Lol I wish I could get spot bonuses from just papers. Paper acceptances don't matter at all.
-

Flan-T5 Performance Analysis: Efficiency Across Language Models
By
–
Pretty cool idea! Great to see Flan-T5 (despite being the smallest model here) hold it's ground pretty well . It even outperforms other LMs like Dolly or StableLM. Also another noteworthy point is that at "compute-match", Flan-T5 3B is equivalent to the cost of a 1.5B
-
Denny Zhou’s 9 Papers Accepted at ICLR 2026
By
–
ICLR should be thrilled to have 9 papers from @denny_zhou
-
Missing Citations in Emergent Abilities Research Paper
By
–
FWIW, i'm not sure how someone can write the mirage paper without citing the canonical emergent abilities paper haha.
-
AI Industry Race: Weekly Model Announcements Growing Token Counts
By
–
twitter these days: announcement: we launched a model trained to 100B tokens!
2nd week announcement: we launch a model trained to 200B tokens
3rd week announcement: we just launch a model at 300B tokens. might as well be twitter-board (not tensorboard). lol. -

Researcher Scooped by Original Paper’s Self-Criticism
By
–
When you think you found a witty rebuttal to a popular paper only to find out that your ideas have been already scooped by the original paper itself. one step ahead bro. Disclaimer: I have not read this fancy "mirage" paper in detail but here's an excerpt from the original
-

RedPajama vs Pile: Dataset Quality Comparison Analysis
By
–
If anything, this graph tells me that RedPajama is worse than the Pile (Pythia) at the same number of pretraining tokens. People can read graphs right? But yea few-shot at 3b is probably really just noise tbh anyway.