it really is incredible what kinds of things become possible when RL on LLMs works. clearly we’re just getting started
@jxmnop
-
Pre-o1 Era: Reflecting on AI Model Evolution
By
–
yep theres a lot to think about here. this was the pre-o1 era ofc
-

Meta Research Internship: Three Months of AI Research Independence
By
–
observations from three months as a Meta research intern in NYC – research jobs are the same everywhere: no one ever asks me what I’m doing or how I’m spending my time
– in particular my recent research has been meandering and potentially going nowhere. but it's nice to have the -

Developer Implements Entire Diffusers Library from Scratch in R
By
–
this absolute maniac implemented the entire diffusers library… from scratch… in R didnt know that was possible (does it run on GPU?). and i dont know if i love it or hate it. but god do i respect it
-

Nano-Transformers: Simplifying Model Implementation Without Dependencies
By
–
huggingface transformers remains extremely useful for getting started but has bloated beyond recognition and is no longer helpful for learning someone should make nano-transformers. implement all the models from scratch without all the dependencies. would be amazing
-

Next Generation GPTs: Hybrid Models Beyond Autoregressive Approaches
By
–
it is increasingly likely that the next generation of GPTs will not be vanilla autoregressive models, nor text diffusion, but some third hybrid thing if you believe this, read subham's research. this paper makes one step in that direction (making diffusion work with kv cache)
-

Mistral’s New Reasoning Model Opens Innovation Opportunities
By
–
this new reasoning model from Mistral looks cool
— dr. jack morris (@jxmnop) 10 juin 2025
it's especially exciting that because no one knows the best way to do RL, there is a lot of room for smaller players to innovate. this wasn't really the case when we spent a few years just making models bigger. https://t.co/Qa5yd5PdZVthis new reasoning model from Mistral looks cool it's especially exciting that because no one knows the best way to do RL, there is a lot of room for smaller players to innovate. this wasn't really the case when we spent a few years just making models bigger.
-
Arithmetic Coding with LLM Probability Distributions Explained
By
–
yes, it’s arithmetic coding using the LLM probability distribution over training token sequences
-
Input Token Proportionality in Language Model Economics
By
–
no…. they are both proportional to the number of input tokens…
-

AI researchers should pursue bigger questions, publish less
By
–
## The case for more ambition i wrote about how AI researchers should ask bigger and simpler questions, and publish fewer papers:
