Starting today our community can test Midjourney V5. It has much higher image quality, more diverse outputs, wider stylistic range, support for seamless textures, wider aspect ratios, better image prompting, wider dynamic range and more. Let’s explore!
MULTIMODAL AI
-
BLIP-2 and Visual ChatGPT: Two Advances in Computer Vision
By
–
BLIP-2:
https://arxiv.org/abs/2301.12597
Visual ChatGPT:
https://arxiv.org/abs/2303.04671 -
BLIP-2: Efficient AI Research Using Frozen Models and Low Compute
By
–
BLIP-2 is a great case study for people who want to do great works with less compute. Amazing how far using frozen models can get you! There are other similar studies such as Visual ChatGPT. VisualChatGPT indeed took it to another level.
-

LLM Achieves 85.4% on US Medical Licensing Exam
By
–
Not everyday a large language model scores 85.4% on the US medical licensing exam (USMLE)
Implications https://
erictopol.substack.com/p/multimodal-a
i-for-medicine-simplified
… -
Carnegie Mellon Multimodal Machine Learning Course Videos
By
–
So good to see courses that are dedicated to this new and vibrant area of AI research. Lecture videos of Multimodal Machine Learning (MML) offered at Carnegie Mellon: https://
youtube.com/playlist?list=
PL-Fhd_vrvisNM7pbbevXKAbT_Xmub37fA
… Follow @Jeande_d for more learning resources and trends in AI research. -
Multimodal Machine Learning: Fusing Vision, Audio, Text, and Actions
By
–
Multimodal machine learning is a hot area in AI research. Unimodal learning has developed massively in the last 5 years. The challenge now is how we fuse different modalities(vision, audio, text, robot actions) into a single agent. GPT-4 & similar models are the beginning.
-

Multimodal Machine Learning Course Carnegie Mellon 2022
By
–
Multimodal Machine Learning – Carnegie Mellon, 2022 A great series of lectures on multimodal machine learning(MML). The course covers fundamental concepts related to MML and recent state-of-the-art MML systems. Lectures: https://
youtube.com/playlist?list=
PL-Fhd_vrvisNM7pbbevXKAbT_Xmub37fA
… Webpage: https://
cmu-multicomp-lab.github.io/mmml-course/fa
ll2022/
… -

VALL-E: Revolutionary Text-to-Speech AI Voice Mimicry Technology
By
–
After ChatGPT and DALL-E, meet VALL-E – the text-to-speech #AI that can mimic anyone’s voice
by @lukekhurst @euronews Read more: https://
buff.ly/3WzDaQh #BigData #MachineLearning cc: @ronald_vanloon @yvesmulkers @pbalakrishnarao -

Metis AI Platform Workshop at Embedded World 2026
By
–
Don't miss our workshop at @embedded_world or online today at 2:30 pm! Join our VP of Product Management, Paul Neil, to learn about our Metis #AI Platform. Register now: https://
talque.com/app#/app/ngx/o
rg/DRl8NY32VVvIa2s0FYTx/session/detail/hQ03E5IfnwJZU86F4c3h
…
#EmbeddedWorld #AI #ComputerVision -

GPT-4 Multimodal Capabilities: Image Input and Recipe Generation
By
–
Now let's get into the details. GPT-4 is multimodal and it now accepts the images as inputs and generates captions, classifications, and analyses. Below is one such example of giving an input image of ingredients and asking GPT-4 to generate a list of recipes.