I don’t think there’s a restriction wrt to multimodal LLMs though (eg using techniques like LLaMA-Adapter v2). But yeah, the evaluation datasets won’t include images.
MULTIMODAL AI
-
What retrieval areas deserve more research focus?
By
–
What areas of retrieval are most interesting to folks and we should push more on?
-
ControlNet Mechanics, Training, and Image Generation Possibilities
By
–
🎥New Videohttps://t.co/xbBLya5pFH
— Satya Mallick (@LearnOpenCV) 21 août 2023
In this video, we'll dissect the mechanics of ControlNet, its training process, and the myriad of image generation possibilities it presents.#ai #computervision #learnopencv pic.twitter.com/ugWLC8IHuQNew https://
youtube.com/watch?v=mac8WF
OvxJQ
… In this video, we'll dissect the mechanics of ControlNet, its training process, and the myriad of image generation possibilities it presents. #ai #computervision #learnopencv -
Real-time GPU rendering and diffusion models performance evolution
By
–
By 2018 or so, on a dedicated GPU, I could get 10-20 fps https://
vimeo.com/382088640#t=58s Current SOTA are all diffusion models, which are multi-step, hence a bit slower again. I bet there will be real-time variants in a year or two. -
Cultural emotional expression data scaling across modalities
By
–
Can I join? Been thinking quite a bit about scaling our cultural emotional expression data, the tricky part has been that cultures use different weightings of face/gesture/voice/context for the same expr. https://
arxiv.org/abs/2103.04262 https://
arxiv.org/abs/2112.05267
v1
… https://
frontiersin.org/articles/10.33
89/fnint.2021.699667/full
… -

AVIS: Autonomous Visual Information Seeking with Large Language Models
By
–
Today on the blog, read all about AVIS — Autonomous Visual Information Seeking with Large Language Models — a novel method that iteratively employs a planner and reasoner to achieve state-of-the-art results on visual information seeking tasks → https://
goo.gle/3P2y2mY -
Baidu Research Advances Computer Vision Robotics with Seven Papers
By
–
Baidu Research RAL is leading the way in innovation! 🚀 Seven innovative papers in #ComputerVision and #Robotics showcase our advances in stereo matching, autonomous excavators, NeRF and more!
— Baidu Research (@BaiduResearch) 18 août 2023
Explore our work: https://t.co/agJCMDfp2q pic.twitter.com/EhFmABG9JOBaidu Research RAL is leading the way in innovation! Seven innovative papers in #ComputerVision and #Robotics showcase our advances in stereo matching, autonomous excavators, NeRF and more! Explore our work: http://
research.baidu.com/Blog/index-vie
w?id=186
… -
Music and Brain Research Inspire AI Innovation
By
–
Both yours and @sciencebanshee's work on music and the brain are so inspiring!
-

ImageBind: Meta AI’s Multimodal Learning Across Six Modalities
By
–
ImageBind by Meta AI is capable of binding information from six modalities. It equips machines with a holistic understanding that connects objects in photos with how they will sound, their 3D shape, temperature & how they move.
— AI at Meta (@AIatMeta) 16 août 2023
More on this research ➡️ https://t.co/yzoIrLRu2p pic.twitter.com/nHOxevSA51ImageBind by Meta AI is capable of binding information from six modalities. It equips machines with a holistic understanding that connects objects in photos with how they will sound, their 3D shape, temperature & how they move. More on this research https://
bit.ly/46hAJaY -

Open Challenges in Large Language Model Research Today
By
–
Open challenges in LLM research The first two challenges, hallucinations and context learning, are probably the most talked about today. I’m the most excited about 3 (multimodality), 5 (new architecture), and 6 (GPU alternatives). Number 5 and number 6, new architectures and