I remember Nicolas teaching this at Cifar with Alex and Ilya there. Those early tries were very important.
@nandodf
-
Tactile Glove Devices: Data Collection Design and Feasibility
By
–
The new efforts to collect touch data with gloves are very exciting, e.g. https://
manus-meta.com and https://
pressureprofile.com/body-pressure-
mapping/tactile-glove
… How hard is it to design and build these devices? It feels that once they become cheap enough, humans (NOT robots) can be paid to collect enough data to -

AI Wearables: Future Adoption and LLM Startup Viability
By
–
The likely future. How long before everyone is using AI assistants via wearables? Does it represent an exit direction for LLM startups desperate to make money? How hard or costly is it to replicate the tech? Would love to hear serious thoughts on this.
-
Generative Video Agents Transform Game Creation and Development
By
–
I highly recommend this talk on generative video agent models for game creation. It is one of the best examples of the blending of video and games that @sama recently tweeted about, and which we should expect to grow in the future.
-
Impressive autonomous vehicle driving demonstration showcased
By
–
I watched this in awe. Insane driving
-
AI Applications Across Industries Drive Job Creation
By
–
I should add that I also really admire people working on applications, eg biology, math, climate, translation, healthcare, history, sustainability, robotics, and so on. It’s nice to see more and more startups too because they help generate much needed new jobs.
-
AI Scaling Safety Efficiency Multimodal Data Collection Priorities
By
–
Revised beliefs: I still think we need to focus on scaling, safety, efficient training and inference, smarter memory (MoEs and long context have addressed this), more modalities (we’re still too slow at collecting visual, sound and touch egocentric data responsibly), online
-
Multimodal AI Leaderboard: Measuring Video, Audio, and Long-Context Performance
By
–
Can someone create a leaderboard with metrics that also measure the features Oriol highlights here: 1. multimodal performance: understanding and generating video, audio, touch, actions, proprioception. 2. long-context: long understanding and generation. I agree Gemini 1.5 Pro
-

Gato: Multimodal LLM Pioneer for Robotics Development
By
–
Fully agree that multimodal LLMs are the solution to robotics. This is why my team pioneered Gato (General AgenT One): https://
arxiv.org/pdf/2205.06175
.pdf
… which we built over two years since 2020 to 2022. One of the most important parts of Gato was the data and its engineering pipeline. -

Multimodal World Models: The Future of Grounded AI Systems
By
–
How to attain Multimodal World Models is a great open question in AI. The solutions will likely lead to more grounded models that interact better with people and make better physics predictions. Hopefully, they will enable scientific generalisation, but this too I feel is an open