Image generation is indeed through the roof…. Everywhere
@nandodf
-
Molmo 72B Model Performance Discussion
By
–
Thanks for suggesting. Molmo 72B is a pretty good model.
-
AI Cambrian Explosion: VLMs, Multimodal Models, and Speech Interfaces
By
–
The Cambrian Explosion in AI this week: numerous VLMs beating SOTA, positive transfer from multimodal to text-only, chain-of-thought takes off, many new speech interfaces, …. more papers than one can read, and more AI people moving than one can keep track of! Whoever thinks
-
Multimodal Models Expected Before AGI Achievement
By
–
@scott_e_reed @yukez I still think multimodal models will come first …. but let's see. Hope you're doing great
-

NVLM Big Models Comparison Results
By
–
Hi Jim, can you add the big models too for a full comparison. I love these results from NVLM:
-
Twitter Experiences AI Cambrian Explosion of Innovation
By
–
Today Twitter feels like the Cambrian Explosion in AI
-

NVLM Paper Reveals Dataset Quality Matters More Than Scale
By
–
The NVLM paper is outstanding. It is full of remarkable findings: (1) "dataset quality and task diversity are more important than scale", (2) positive transfer from multimodal datasets to text-only on math benchmarks, (3) model and data ablations, etc. Congrats to the authors on
-
OpenAI o1 Models Advance AI Research at Microsoft
By
–
The @OpenAI o1 models represent one of the smartest advances in AI in a long time. Having just joined @Microsoft AI, one of the things I really look forward to is being able to contribute to some of these fruitful ideas to advance OpenAI’s mission. The opportunity to work
-

Anil Seth on Consciousness, Perception and Intelligence Research
By
–
I had the fortune of listening to
@anilkseth
today at Microsoft London on consciousness, perception, intelligence, life and diversity of human minds. Fascinating. Highly recommend his works. I’m glad that there are smart people like him exploring these research topics. -
Video Generation AI Models: Benchmarks and Training Data Comparison
By
–
Great video generation. I wish we had good public benchmarks to compare to other approaches like Veo, Runway Gen-3, Sora, Haiper, etc. It would also be nice to know what data is being used for (pre-)training in each case, and to know how influential data and model scale are.… https://t.co/xbIXAk9q2L
— Nando de Freitas (@NandoDF) 19 septembre 2024Great video generation. I wish we had good public benchmarks to compare to other approaches like Veo, Runway Gen-3, Sora, Haiper, etc. It would also be nice to know what data is being used for (pre-)training in each case, and to know how influential data and model scale are.