Runway Aleph is now available to all Enterprise accounts and Creative Partners. Edit, transform and generate videos like never before across endless new use cases. To get started, enter Chat Mode or find our guide at the link below.
MULTIMODAL AI
-
Brain Vision Processing Gap in Modern LLMs Architecture
By
–
the human brain reserves 40% of its processing exclusively for vision. modern LLMs somehow evolved without this entirely
-

MMBench-GUI: Multi-Platform Evaluation Framework for GUI Agents
By
–
MMBench-GUI Hierarchical Multi-Platform Evaluation Framework for GUI Agents
-

VLMs Outperform with Visual Data Alone Over Numerical
By
–
Interestingly, VLMs perform better with plot alone compared to plot + data! This paper shows that GPT‑4.1 and Claude 3.5 excels at scatterplot analysis, and could often perform better only by looking at the plot, without reading the actual numerical data.
-
AI Image Generation Quality: Pelican on Bicycle Benchmark
By
–
I'll change my benchmark when a model produces a genuinely good picture of a pelican riding a bicycle! These bicycles see still pretty obviously bad
-

Gemini 2.5 Conversational Segmentation: Natural Language Vision AI
By
–
🖼️ Conversational Segmentation (Gemini 2.5): hype or helpful?
— Louis-François Bouchard 🎥🤖 (@Whats_AI) 28 juillet 2025
• Segments with natural‑language prompts (vs Meta SAM’s points/clicks/boxes/masks).
• Handles relationships, conditionals, OCR text, and multilingual labels.
• Great for prototyping and LLM apps; edge/real‑time… pic.twitter.com/APadBDIMn6Conversational Segmentation (Gemini 2.5): hype or helpful? • Segments with natural‑language prompts (vs Meta SAM’s points/clicks/boxes/masks).
• Handles relationships, conditionals, OCR text, and multilingual labels.
• Great for prototyping and LLM apps; edge/real‑time -
GLM-4.5 generates animated SVG of pelican on bicycle
By
–
GLM-4.5 vibe coding
— AK (@_akhaliq) 28 juillet 2025
Generate a animated svg of a pelican riding a bicycle pic.twitter.com/Se4xkNhS2KGLM-4.5 vibe coding Generate a animated svg of a pelican riding a bicycle
-
Image pretraining impact on text-based AGI evaluation benchmarks
By
–
maybe our crux here is that most of the 'AGI evals' are text-based. so my point is that adding image pretraining doesn't help on any of the (text-based) evals.
-
Smart AI Suggestions Help Users Move Forward Contextually
By
–
Also coming – smart suggestions based on what you're doing, to help you move forward. That could be a video tutorial after you've read 4 sweater how-to's, or trusted sites to get ordained if you've been researching tips for officiating a wedding.
-
Google Launches Flow: New AI Video Creation Tool
By
–
Google has a new tool just for making AI videos (Flow) #AI #AIio #AIInnovation #ML #DataScience #Futureofwork @HaroldSinnott @fogoros @iainljbrown @NandoDF @katecrawford @drhassanrashidi @YuHelenYu