GitHub – guoyww/AnimateDiff: Official implementation of AnimateDiff. https://
bit.ly/44Ca6w0
Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning #AI #MachineLearning #DeepLearning #LLMs #DataScience
MULTIMODAL AI
-

AnimateDiff: Animate Text-to-Image Diffusion Models Without Tuning
By
–
-
Google Robots Become Smarter With AI Language Models
By
–
Aided by #AI Large Language Models, Google's Robots Are Getting Smart https://
nytimes.com/2023/07/28/tec
hnology/google-robots-ai.html
… -
Multiple Generated People Tend to Look Similar
By
–
Also, when you generate multiple people, they all tend to look similar.
-

Google’s Bard AI Now Supports Voice and Visual Recognition
By
–
Now Google’s Bard AI chatbot can talk and respond to visual prompts https://
bit.ly/3OoklhV #AI #MachineLearning #DeepLearning #LLMs #DataScience -
PaLI-X advances robotics with real-world apple and banana sorting
By
–
This is cool! 🥳 Nice to see more impact from PaLI-X (ViT22B + 32B UL2) on robotics.
— Yi Tay (@YiTayML) 29 juillet 2023
Gotta admit the ability to move apples and bananas in the real world just hits different from any emergent language/vision tasks. https://t.co/izPwjiVEoMThis is cool! Nice to see more impact from PaLI-X (ViT22B + 32B UL2) on robotics. Gotta admit the ability to move apples and bananas in the real world just hits different from any emergent language/vision tasks.
-
Robot Motions as Language: Transformer Architecture for Action Understanding
By
–
Interesting idea: treat sequences of robot motions as sequences of words with a similar transformer architecture to link language, vision, and action.
-

RT-2 Robotics Transformer: Vision and Language to Action Model
By
–
RT-2 relevant links: Project page: https://
robotics-transformer2.github.io
Paper: https://
robotics-transformer2.github.io/assets/rt2.pdf
Blog: https://
deepmind.com/blog/rt-2-new-
model-translates-vision-and-language-into-action
… Other links:
RT-1: https://
robotics-transformer1.github.io
PaLI-X(direct scale up of PaLI): https://
arxiv.org/abs/2305.18565
PaLM-E: https://
arxiv.org/abs/2303.03378 Work from Google -

RT-2: Vision-Language-Action Models for Robotic Control
By
–
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control RT-2(Robotic Transformer 2) is a new model in family of vision-language action models(VLAs). RT-2 builds up on RT-1 and vision-language models(VLMs). RT-2 makes use of pre-trained VLMs and is
-
LLM Overkill: Do We Really Need AI for Robot Vision Tasks?
By
–
while cool, I’m not sure you need an LLM to identify an extinct animal in an image, translate it’s coordinates to a robot arm that can then pick it up.