Curious to know more about DePlot, the tool for multi-modal chain-of-thought reasoning on plots and charts? Drop by the #ACL2023 Google booth today at 3:30pm where @eisenjulian will explain how it is able to answer complex questions about charts and even summarize content!
MULTIMODAL AI
-

Google unveils Gemini amid AI disorganization
By
–
Google has announced a new AI model from DeepMind: Gemini. Yet another model, yet another name to compete with 'ChatGPT' after LaMDA, Bard, and others… All of this is just proof of the complete disorganization within Google’s AI efforts. 150 teams working on it.
-

Generalization Gap in Visual Robotic Imitation Learning
By
–
Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation paper page: https://
huggingface.co/papers/2307.03
659
… What makes generalization hard for imitation learning in visual robotic manipulation? This question is difficult to approach at face value, but the -

GPT4RoI: Instruction Tuning LLM on Region-of-Interest
By
–
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest paper page: https://
huggingface.co/papers/2307.03
601
… Instruction tuning large language model (LLM) on image-text pairs has achieved unprecedented vision-language multimodal abilities. However, their vision-language -

Computer Vision Newsletter #33: Artificial Intelligence and Deep Learning
By
–
Computer Vision Newsletter #33
https://bit.ly/3BImZrO #AI #MachineLearning #DeepLearning #LLMs #DataScience -
Bookmarking Key Papers on Learning and Vision-Language Models
By
–
@chelseabfinn on learning to learn with gradients. I bookmarked it and hope to read it sometime: https://
ai.stanford.edu/~cbfinn/_files
/dissertation.pdf
… Also @karpathy on connecting images and texts, which was ahead of time given current progress in visual language learning: https://
cs.stanford.edu/people/karpath
y/main.pdf
… -
AR and AI Transform Second-Hand Sales Applications Forever
By
–
El futuro de las aplicaciones de venta de segunda mano está a punto de cambiar para siempre. Gracias al uso de la realidad aumentada y a la inteligencia artificial, el proceso de poner a la venta un producto que no utilizamos se simplifica enormemente.
— Juan Merodio (@juanmerodio) 8 juillet 2023
En este proyecto… pic.twitter.com/k4r6vcplXXEl futuro de las aplicaciones de venta de segunda mano está a punto de cambiar para siempre. Gracias al uso de la realidad aumentada y a la inteligencia artificial, el proceso de poner a la venta un producto que no utilizamos se simplifica enormemente. En este proyecto
-

AI Model Decodes Thoughts into Images with 80% Accuracy
By
–
Two researchers have created a new #AI model that can draw what you’re thinking with 80% accuracy https://
bit.ly/3FTXOoD -
Emotional AI Recognition: Low-Level Contagion vs. Theory of Mind
By
–
Oooh. Strikes me that one use the same method while having people view emotional images, some where it's low level emotional contagion, and others where it's high level theory of mind emotion recognition? Or caption an image on its objective content vs. emotional content…
-
CodeVQA: Python Code Framework for Visual Question Answering
By
–
Introducing CodeVQA, a framework that generates Python code, paired with simple visual functions, to enable image processing for few-shot visual question answering. Learn how CodeVQA performs multi-step visual reasoning and outperforms prior work → https://t.co/LQaZikRShe pic.twitter.com/1qcPWmkq1V
— Google AI (@GoogleAI) 7 juillet 2023Introducing CodeVQA, a framework that generates Python code, paired with simple visual functions, to enable image processing for few-shot visual question answering. Learn how CodeVQA performs multi-step visual reasoning and outperforms prior work → https://
goo.gle/46w6Zam
