Similar to how text-based LLMs were made possible via large pre-trained text transformers, the progress on large pre-trained vision transformers has been swift. This includes Meta's fantastic work (XCiT, DINO, DINOv2, SAM), Landing AI's work on Visual Prompting, and the work of
MULTIMODAL AI
-

Multi-Modal Classifiers for Open-Vocabulary Object Detection
By
–
Multi-Modal Classifiers for Open-Vocabulary Object Detection paper page: https://
huggingface.co/papers/2306.05
493
… The goal of this paper is open-vocabulary object detection (OVOD) x2013 building a model that can detect objects beyond the set of categories seen at training, thus enabling the -

Social Impact Evaluation of Generative AI Systems Across Modalities
By
–
Evaluating the Social Impact of Generative AI Systems in Systems and Society paper page: https://
huggingface.co/papers/2306.05
949
… Generative AI systems across modalities, ranging from text, image, audio, and video, have broad social impacts, but there exists no official standard for means of -

AI Model Draws Thoughts with 80 Percent Accuracy
By
–
Two researchers have created a new #AI model that can draw what you’re thinking with 80% accuracy https://
bit.ly/3FTXOoD -
Open-source multimodal models ahead of GPT-4
By
–
Second, there are already several multimodal models out there; GPT-4 is still a pure text model. Here’s where open-source innovated first.
-
Multimodal LLMs and reasoning advancement in foundation models
By
–
Good points from @ylecun albeit it may just be a matter of time before we have foundation AI LLM models trained on + taking inputs on a multimodal basis (including video and text). However, logic and reasoning whilst improving with Chain of Thought may still take longer to
-
KaiberAI: AI-Powered Video Processing Tool Explained
By
–
It’s a real video recording processed by a tool called @KaiberAI
-
ControlNet Pose Technique Achieves Impressive Results
By
–
+1 looks awesome. Controlnet pose just hits different,
-

Video-LLaMA: Instruction-tuned Audio-Visual Language Model
By
–
-LLaMA: An Instruction-tuned Audio-Visual Language Model for Understanding Hang Zhang, Xin Li, Lidong Bing: https://
arxiv.org/abs/2306.02858 #DeepLearning #LargeLanguageModels #LLM -

REMEDIS: Robust Medical Imaging with Representation Learning
By
–
This is neat: “a representation-learning strategy for ML models applied to medical-imaging tasks that mitigates such ‘out of distribution’ performance problem and that improves model robustness and training efficiency.” “REMEDIS (for ‘Robust and Efficient Medical Imaging with