CulturalGround data construction: We designed a scalable pipeline to create culturally grounded multilingual VQA data from Wikidata, a structured knowledge base. Our data curation pipeline:
– Cultural Entity Selection: Extract 3M+ culturally relevant entities from Wikidata
MULTIMODAL AI
-

CulturalGround: Scalable Multilingual VQA Data Pipeline from Wikidata
By
–
-
Embodied Intelligence and AI Systems Interacting with Physical World
By
–
Ask us your questions about embodied intelligence or AI systems that interact w/the world. We’re featuring a few in an upcoming explainer w/MIT prof. Vincent Sitzmann (
@vincesitzmann
). For more on his work: https://
vincentsitzmann.com -

Creating Videos from Single Images with Grok Technology
By
–
Created this video using a single image using grok. Quite impressive https://t.co/BUHvaD48n5 pic.twitter.com/bYRKNfOelj
— Satya Mallick (@LearnOpenCV) 12 août 2025Created this video using a single image using grok. Quite impressive
-
Pixel Reasoner: Open-Source VLM Reasoning Framework Breakthrough
By
–
🚀 New Breakthrough in Vision-Language Reasoning 🧠🖼️
— Dr. Debashis Dutta (@debashis_dutta) 12 août 2025
Originally shared by @WenhuChen — Pixel Reasoner is the first open-source framework enabling Vision-Language Models (VLMs) to “think in pixel space” through curiosity-driven reinforcement learning.
💡 The Challenge
Most… pic.twitter.com/IXCFVuNMeMNew Breakthrough in Vision-Language Reasoning Originally shared by @WenhuChen — Pixel Reasoner is the first open-source framework enabling Vision-Language Models (VLMs) to “think in pixel space” through curiosity-driven reinforcement learning. The Challenge Most
-
Audio-Driven AI Model Supports Multiple Languages
By
–
yes, it's an audio driven model, feel free to upload audios in different languages
-
Why Frontier VLMs Lag Behind LLMs and Vision Models
By
–
Frontier LLM have superhuman text-based world knowledge. Frontier image / video models have superhuman vision-based world knowledge (e.g. Genie). But current frontier VLMs are still absolutely clown shoes. Why? Relative scarcity of image:text pairs (while there is plenty of text
-
Video Generation AI Now Available Web iOS Android
By
–
Bring ideas to life with video generation, now available on web, iOS and Android.
— Perplexity (@perplexity_ai) 11 août 2025
Pro subscribers can create 5 videos/month, Max can generate 15/month with enhanced quality.
Ask, create, inspire. Ideas are better when you can see them. pic.twitter.com/BMYcgEQvIBBring ideas to life with video generation, now available on web, iOS and Android. Pro subscribers can create 5 videos/month, Max can generate 15/month with enhanced quality. Ask, create, inspire. Ideas are better when you can see them.
-

GLM-4.5V: Open-source vision model dominates benchmarks, exciting open-source community
By
–
GLM-4.5V is a new open-source vision model (just released) and it's dominating benchmarks. Glad to see serious teams prioritizing open-source models – these results are wild: Hugging Face: http://
huggingface.co/zai-org/GLM-4.
5V
…
GitHub: http://
github.com/zai-org/GLM-V http://
Z.ai API: -

Google’s Genie 3 Physics Simulation: Plane Surface Collision Behavior
By
–
Googles Genie 3 is insane. Look at the physics. The plane is bumping off when hitting the surface. pic.twitter.com/tVI2oGoc0x https://t.co/YETfXfu981
— Chubby♨️ (@kimmonismus) 11 août 2025Googles Genie 3 is insane. Look at the physics. The plane is bumping off when hitting the surface.
-
Meta FAIR Wins Algonauts 2025 with TRIBE Brain Modeling
By
–
🏆 We're thrilled to announce that Meta FAIR’s Brain & AI team won 1st place at the prestigious Algonauts 2025 brain modeling competition.
— AI at Meta (@AIatMeta) 11 août 2025
Their 1B parameter model, TRIBE (Trimodal Brain Encoder), is the first deep neural network trained to predict brain responses to stimuli… pic.twitter.com/IeX5gPd8GzWe're thrilled to announce that Meta FAIR’s Brain & AI team won 1st place at the prestigious Algonauts 2025 brain modeling competition. Their 1B parameter model, TRIBE (Trimodal Brain Encoder), is the first deep neural network trained to predict brain responses to stimuli