Molmo and PixMo Open Weights and Open Data for State-of-the-Art Multimodal Models discuss: https://
huggingface.co/papers/2409.17
146
… Today's most advanced multimodal models remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good
GENERATIVE AI
-

Molmo and PixMo: Open-Weight Multimodal Models with Open Data
By
–
-

DreamWaltz-G: Expressive 3D Gaussian Avatars from 2D Diffusion
By
–
DreamWaltz-G
— AK (@_akhaliq) 26 septembre 2024
Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion
discuss: https://t.co/szvheoe5MM
Leveraging pretrained 2D diffusion models and score distillation sampling (SDS), recent methods have shown promising results for text-to-3D avatar generation.… pic.twitter.com/BP9ckpAZFPDreamWaltz-G Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion discuss: https://
huggingface.co/papers/2409.17
145
… Leveraging pretrained 2D diffusion models and score distillation sampling (SDS), recent methods have shown promising results for text-to-3D avatar generation. -

TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans
By
–
TalkinNeRF
— AK (@_akhaliq) 26 septembre 2024
Animatable Neural Fields for Full-Body Talking Humans
discuss: https://t.co/ous6cqa59g
We introduce a novel framework that learns a dynamic neural radiance field (NeRF) for full-body talking humans from monocular videos. Prior work represents only the body pose or… pic.twitter.com/vnFyv4F3eJTalkinNeRF Animatable Neural Fields for Full-Body Talking Humans discuss: https://
huggingface.co/papers/2409.16
666
… We introduce a novel framework that learns a dynamic neural radiance field (NeRF) for full-body talking humans from monocular videos. Prior work represents only the body pose or -
OpenAI Eyes For-Profit Status, Equity Stake for Sam Altman
By
–
not only is openai considering restructuring to become a for-profit business, but it's also discussing giving sam altman a 7% equity stake. read the latest from me, @dinabass
, @shiringhaffary
, and @EdLudlow
. -
Multimodal AI Models Text-Only Processing Capabilities
By
–
The reason I ask is that generally for a multimodal model I want to be able to do some text-only stuff too.
-
Multimodal AI Model Handles Text and Coding Tasks
By
–
Yup that's what I mean. I tried doing some text-only chat on the playground but it looks like it actually requires an image input. So I put in a random pic and asked it coding questions, and it seemed to do ok.
-
11B and 90B Models Now Support Multimodal Tasks and Text Function Calls
By
–
The 11B and 90B currently support multimodal tasks and text-only function calls.
-
Google Releases Improved Gemini 1.5 Flash and Pro Models
By
–
The new improved Gemini 1.5 Flash and Pro production models are the best models by far on a performance / price basis. You should try them for all your workloads at scale.
-
New 1B and 3B Models Available for Mobile Phones
By
–
We can't wait to hear how you feel about running the new 1B & 3B models on your phone.