Very cool! Would be awesome if you ended up sharing some models or datasets. Feels like we're close to getting to real breakthroughs in 3D
MULTIMODAL AI
-
VideoReTalking Model Achieves Impressive Lip-Sync Quality
By
–
I'm blown away by the lip-sync quality of the VideoReTalking model.
— fofr (@fofrAI) 3 novembre 2023
– created an image with my SDXL fine-tune
– passed img to @runwayml Gen2 to animate
– used video and VideoReTalking model to lip-sync:https://t.co/MAzbb5Beujhttps://t.co/TXK40Ck4Qf https://t.co/bQdy46DtVd pic.twitter.com/h1T0OpcTKsI'm blown away by the lip-sync quality of the VideoReTalking model. – created an image with my SDXL fine-tune
– passed img to @runwayml Gen2 to animate
– used video and VideoReTalking model to lip-sync: https://
replicate.com/cjwbw/video-re
talking
… https://
replicate.com/p/ql7ay5lbno2i
s3no2wpwvzsrna
… -

SDXL Fine-Tuned on Toy Story Characters Generates Quirky Images
By
–
I trained SDXL on the slightly weird looking people from Toy Story (1995). Inspired by @skirano
’s really good Toy Story tune. Just more weird. https://
replicate.com/fofr/sdxl-toy-
story-people
… -
Multimodal AI, Synthetic Data, and Enterprise AI Readiness Trends
By
–
I could talk for hours about the future of AI, but right now, I’m really looking at multimodal AI, synthetic data, companion AI, GPU shortage mitigation, hyperpersonalization, and AI enterprise readiness trends (from training to LLM ops to quick deploy templates).
-
DALL-E 3 and Fine-tuning Capabilities Comparison
By
–
Yeah, Dalle3 is amazing too, it’s different use-cases too, /tune feels a lot like a “simplified” fine-tuning.
-

Midjourney Style Tuner Creates Amazing Community Art in 24 Hours
By
–
Style Tuner – 24 hours and the Midjourney community is going crazy! Here are some of the most incredible things the community members have created so far. Here are the top finds + tutorials from each creator:
-
AI Image Captioning Tool for Emotional Salient Features Analysis
By
–
We also offer a captioning method to convert images into sentences describing emotionally salient features related to face, body, pose, human interactions, activity and environment. Social scientists might be interested in our manual annotation interface
-
Emotional AI: Fast Slow Emotion Recognition Vision Models
By
–
Captioning interface
Arxiv: https://
arxiv.org/abs/2309.13136
Project: https://
rosielab.github.io/emotion-captio
ns/
… Emotions Fast and Slow Arxiv: https://
arxiv.org/abs/2310.19995
Project: https://
rosielab.github.io/Fast-and-Slow/ -

CLIP and LLaVA Struggle with Contextual Image Interpretation
By
–
In this example, we believe CLIP sees the "surprised cat pose" and predicts doubt, surprise and fear, ignoring context. Oddly, LLaVA has also gone a bit far, inferring the person in this image is experiencing sadness because it's their last ski trip of the season.
-
Multimodal AI Shows Promise for Emotion Recognition Beyond Zero-Shot Methods
By
–
Overall, these zero shot methods do not perform nearly as well humans nor models tuned on the training data from EMOTIC (Emotions in Context). But there appears to be promise in linking vision with language to account for both affective and cognitive empathy