Let's start with one of the mindblowing use cases: Here is an example of GPT4 turning a picture of a sketch into a fully functioning website:
MULTIMODAL AI
-
Realtime Emotion Detection AI: Can Technology Truly Capture Human Complexity?
By
–
Realtime emotion detection with #AI is an impressive technological advancement, but can it truly capture the complexity of human emotions?
— Pascal Bornet (@pascal_bornet) 15 mars 2023
What do you think?
Credit: Rondinelli Morais#innovation #artificialintelligence #machinelearning pic.twitter.com/Xl14BPY73VRealtime emotion detection with #AI is an impressive technological advancement, but can it truly capture the complexity of human emotions? What do you think? Credit: Rondinelli Morais
#innovation #artificialintelligence #machinelearning -
Microsoft Visual ChatGPT: Image Understanding and Generation
By
–
Microsoft’s Visual ChatGPT Enables Image Understanding and Generation https://
syncedreview.com/2023/03/14/mic
rosofts-visual-chatgpt-enables-image-understanding-and-generation/
… -
GPT-4 to interpret images and documents, announced by @aibreakfast
By
–
GPT-4 will have the ability to interpret images and documents pic.twitter.com/cEJFAzxyni
— AI Breakfast (@AiBreakfast) 14 mars 2023GPT-4 will have the ability to interpret images and documents
-
GPT-4 Vision and Language Capabilities with Visual Input Examples
By
–
GPT-4 can take two important modalities that humans rely on – vision and language. Its output is text though(maybe in the future it will be many-to-many). GPT-4 can explain images, helps you understand what's in images, decode memes, etc. A few examples of GPT-4 on visual input.
-

GPT-4 Achieves Human-Level Performance on Academic Benchmarks
By
–
Some interesting things about GPT-4 GPT-4 exhibits human-level performance on a range of language and vision academic benchmarks and shows excellent performance in professional exams. Below is GPT-4 performance compared to GPT-3.5 & GPT-4(no vision)>> vision improves language!?
-
GPT-4 Released: Multimodal AI with Vision Capabilities
By
–
GPT-4 is finally out. As many people alluded to lately, GPT-4 is multimodal. It can take images and texts. It can answer questions about images, converse back & forth, etc. GPT-4 is essentially better GPT-3.5 + vision. Blog: https://
openai.com/research/gpt-4
Paper: https://
cdn.openai.com/papers/gpt-4.p
df
… -
LinkedIn text-to-video feature: Where is the promised capability?
By
–
Where is the text2video LinkedIn promised me?
-
Vision Capability Rollout Excitement and Uncertainty
By
–
Plot twist: it's solved or probably it's not solved or we're not sure . Really looking forward the vision capability rolling out publicly soon, unlocks a ton of new/exciting uses.




