Mono-InternVL-1.5 Towards Cheaper and Faster Monolithic Multimodal Large Language Models
MULTIMODAL AI
-

Nested Matryoshka Clustering for Scalable Visual Representation Learning
By
–
Franca Nested Matryoshka Clustering for Scalable Visual Representation Learning
-

Rodney Brooks Reflects on 2006 AI Prediction Benchmarks
By
–
You asked about how my predictions are holding up. My last slide at Dartmouth '06 was 4 immediate challenges: 1. Visual object recognition capabilities of a 2 year old child, 2. Manual dexterity of a 6 yr old, 3. Language capabilities of a 4 yr old, 4. social sophistication of a
-
AI Music Generation Brings Rilke’s Poetry to Life with Emotion
By
–
It is has not ceased to be weird that I can put Rilke’s First Elegy into Suno and get out a coherent 8 minute performance with music.
— Ethan Mollick (@emollick) 21 juillet 2025
You might not like the interpretation, but it is genuinely amazing that this audio, with apparent emotion, is all 100% AI from the verses alone. pic.twitter.com/6WbkpO4aB0It is has not ceased to be weird that I can put Rilke’s First Elegy into Suno and get out a coherent 8 minute performance with music. You might not like the interpretation, but it is genuinely amazing that this audio, with apparent emotion, is all 100% AI from the verses alone.
-
Multimodal AI: The Next Frontier Reshaping Technology Interaction
By
–
The Next AI Frontier – How Multimodal Systems Are Reshaping Our World From vision to voice and beyond, multimodal AI is the next leap in intelligent systems. Find out how it’s changing the way we interact with technology. Read more https://
bernardmarr.com/the-next-ai-fr
ontier-how-multimodal-systems-are-reshaping-our-world/
… #MultimodalAI -
MirageLSD Brings Real-Time Video Diffusion Models to Life
By
–
Les modèles IA de diffusion vidéo, mais… en temps réel !
— VISION IA (@vision_ia) 20 juillet 2025
Les filtres vidéo classiques sont en temps réel, mais limités à des effets basiques (recolorisation, style simple).
Les modèles de diffusion (comme Veo) sont bien plus puissants, mais lents.
Voici MirageLSD, c’est la… https://t.co/fPzR5copjjLes modèles IA de diffusion vidéo, mais… en temps réel ! Les filtres vidéo classiques sont en temps réel, mais limités à des effets basiques (recolorisation, style simple). Les modèles de diffusion (comme Veo) sont bien plus puissants, mais lents. Voici MirageLSD, c’est la
-
AI-Generated Video with Google Veo 3 and Higgsfield Soul ID
By
–
Tout cela est généré par l’IA.
— VISION IA (@vision_ia) 20 juillet 2025
Son visage, son corps, sa voix, ses vêtements, les mouvements de la caméra, le monde autour d’elle… tout est modifiable.
Créé avec Google Veo 3 et Higgsfield Soul ID.
Des exemples : pic.twitter.com/Uw3BHokRTiTout cela est généré par l’IA. Son visage, son corps, sa voix, ses vêtements, les mouvements de la caméra, le monde autour d’elle… tout est modifiable. Créé avec Google Veo 3 et Higgsfield Soul ID. Des exemples :
-
AI generative systems democratize creative production across modalities
By
–
These are systems that respond to human writing and (often) techniques that apply to human psychology. Everyone now has a machine that makes words, images, video, sound where the limit is often your own ability to imagine something new (or invoke old ideas others do not know).
-

Imagen 4 LoRA Generalizes Well to Animal Fur Rendering
By
–
Interestingly it seems to have also generalised well to fur on animals too. Here's an Imagen 4 cat on the left, and the lora on the right
-

Imagen 4 LoRA Outperforms in Facial Hair and Details
By
–
Does guys too. Imagen 4 on left, lora on right. It's doing very well on beards and stubble. One of the bugs with this lora is that it likes to give people ear piercings, still need to fix that.
