New paper- A few months ago, Meta released Segment Anything Model (SAM). It's already in photography apps, medicine, and video-generation. Now the Team invented SAM’s little brother, EfficientSAM. It is small but mighty! ****
With 20x fewer parameters and 20x faster
MULTIMODAL AI
-

Meta’s EfficientSAM: 20x Smaller, 20x Faster Vision Model
By
–
-

Googly Eyes SDXL Fine-Tune: Add Playful Eyes to Any Image
By
–
Googly-eyes SDXL fine-tune. Add googly eyes to anything. https://
replicate.com/fofr/sdxl-goog
ly-eyes
… Side effect – it usually makes a face out of it -
Pika 1.0 Expand Canvas Transforms 16:9 Videos Into Anamorphic Masterpieces
By
–
🌟Turn your 16:9 videos into anamorphic masterpieces with the Pika 1.0 'Expand Canvas' feature. 🌟
— Pika (@pika_labs) 9 décembre 2023
Sign up at https://t.co/nqzjGy82Lx pic.twitter.com/5Oxq2j4armTurn your 16:9 videos into anamorphic masterpieces with the Pika 1.0 'Expand Canvas' feature. Sign up at http://
pika.art -

Video-LLaVA: Unified Visual Representation Learning for Multimodal AI
By
–
GitHub – PKU-YuanGroup/Video-LLaVA: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection https://
bit.ly/3Rab0KS #AI #MachineLearning #DeepLearning #LLMs #DataScience -
AI and Art: Exploring Creative Possibilities with Refik Anadol
By
–
Join Mira Lane, Senior Director of Technology & Society at Google, and renowned media artist @RefikAnadol as they explore the power of #AI in art in this new Dialogues episode. Full video →https://t.co/nSTdRMQgKg
— Google AI (@GoogleAI) 8 décembre 2023
More from the Dialogues series →https://t.co/ESdHHZ2S08 pic.twitter.com/6CZbDl8gWlJoin Mira Lane, Senior Director of Technology & Society at Google, and renowned media artist @RefikAnadol as they explore the power of #AI in art in this new Dialogues episode. Full video →
https://
goo.gle/3GnS3iO More from the Dialogues series →
https://
goo.gle/3t5iGpH -
AI Misrepresentation: Video Understanding Claims Under Scrutiny
By
–
They gave the impression that the AI was responding to the video footage. I can understand the speeding up responses. But the implication that it was responding to the video footage it was seeing when it wasn't even close to doing that, feels wrong.
-
Exciting new LLM interfaces and multimodal applications emerging
By
–
I am particularly excited to see lots of cool new interfaces for LLM’s along with people pushing the multi-modal use cases to the limit. There’s so much to build!
-

Generate Every Emotion in Any Style with Gen-2
By
–
Generate every emotion, in any style, with Gen-2.https://t.co/ekldoIshdw pic.twitter.com/J5QqR93y6t
— Runway (@runwayml) 8 décembre 2023Generate every emotion, in any style, with Gen-2. http://
runwayml.com -
Exponential difficulty: AI accuracy improvements beyond human baseline
By
–
Once you get pretty accurate, each additional percentage point of accuracy is both exponentially harder to get, and qualitatively different from the human standpoint. Don’t underestimate what they have been able to accomplish yet.
-
Concept Association Bias in Vision-Language Models: Contrastive vs Autoregressive
By
–
We call this Concept Association Bias (CAB). We’ve found that models trained using contrastive loss (e.g. parts of BLIP and BLIP-2) also have CAB. However, models trained with autoregressive loss (e.g. OFA and BLIP-2-FlanT5) don't exhibit this bias. 4/5