From 2011: "Detecting temporal patterns and predicting into the future is a fundamental problem in machine learning. It has gained great interest recently in the areas of nonparametric Bayesian statistics (Wood et al., 2009) and deep learning (Sutskever et al., 2011), with
@nandodf
-
Lambda values and encoder-decoder architecture in research execution
By
–
Happy to see this execution of a very interesting research direction. Question: did you try more extreme values of lamda? Could the lack of obvious trend in this (fig 3) be the result of variance? What would be its likely source? Also, did you consider using an encoder-decoder
-
Credit Assignment and Multi-Step RL in LLM Research
By
–
Credit assignment is hard. I wonder how many LLM papers use multi-step RL? Tool use is the thing that comes to mind. It would be great if someone working on this could comment. Also, how many people out there are doing multi-step RL with LLMs?
-
Acknowledging sources and adding personal reasoning analysis
By
–
This is the source … of course based on reading other people’s stuff and adding a bit of reasoning on top
-
Scaling trends in latest AI models and their development
By
–
They seem to scale, eg o1, R1, Qwen3, MAI1 …
-
Questions on masked autoencoders, latent diffusion, and rollout loss consistency
By
–
This was very instructive. Thank you. Two questions: (1) have people replaced the masked autoencoder with latent diffusion? (2) in the rollout loss why not use multiple rollouts and consistency checks as in one of our old papers on learning awareness models
-
Traditional RL Researchers and Their Opposition to Supervised Learning
By
–
Traditional RL folks were for some weird reason against supervised learning and LLMs. Have they learned the bitter lesson?
-
MAI1 preview: New young team behind major AI breakthroughs
By
–
Thanks for asking: MAI1-preview, MAI voice. Most of the team is less than a year old, O(100) and made of the people that just recently helped ship MovieGen, Imagen3, Veo, Gemini, Llama, NotebookLM, Genie, etc, …. Like the 90s La-di-da, da-di-da
Di-dai-dai-da
More than words -
Young team achieves rapid progress in LLM and voice technology
By
–
Maybe not so. The team has existed for less than a year and has already landed a LLM preview in lm arena and several voice products – I believe the most expressive, efficient and highest quality voice generation in history. We’re improving fast, and whoever applies will make many
-

Standard Maximum Likelihood Expectation Maximisation EM Interpretation
By
–
In fact, this may be interpreted as the standard maximum likelihood expectation maximisation (EM).