AI NEWS: Google just added AI avatars into its Vids AI video editing platform. Plus, more news from OpenAI, Tencent, Nous Research, Microsoft, Moonshot AI, and Prime Intellect. Here's everything you need to know:
MULTIMODAL AI
-
Google Releases Gemini 2.5 Flash Multimodal AI Model
By
–
Gemini 2.5 flash image preview aka nano-banana
-

AI Image Generation Detail Consistency in Transformation Output
By
–
Cannot believe the level of detail consistency in this AI-generated transformation. Taylor's and Travis' faces change a bit, but take a second to notice the smaller details in this Nano Banana image gen output. Notice the now visible Eagle's logo on Jason's shoulder, the glove
-

LLM Limitations and Rapid Progress in Image Capabilities
By
–
I agree that it is a problem that the models have no idea of their own limits, it is one of many issues that make LLMs hard to use. And yes, agree image comprehension and image creation are both limited, but the evidence suggests pretty rapid improvement & some real utility.
-
AI Vision Models: Weaknesses in Counting and Image Generation
By
–
Clear weak spots remain counting, generating alternate images when the training data is thick (full glasses of wine, clocks with oddly set hands), etc. It isn't hard to make them fail. But there is a lot they do very well, and the gains have been pretty quick so far.
-

Image Generation Progress: Spaghetti Forks and LLM Limitations
By
–
Well, I got six forks made of spaghetti on the first try, but one is a double-sided fork It is pretty amazing how far imagegen has come in the past years (they aren't flawless, but this would have been impossible months ago). Yet they aren't really a good measure of LLM ability
-
MiniCPM-V 4.5 Chat App Development with Impressive Benchmark Performance
By
–
vibe coding a MiniCPM-V 4.5 @OpenBMB chat app in anycoder
— AK (@_akhaliq) 28 août 2025
MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro,… pic.twitter.com/r1i5b6JfFpvibe coding a MiniCPM-V 4.5 @OpenBMB chat app in anycoder MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro,
-

Self-Rewarding Vision-Language Model Through Reasoning Decomposition
By
–
Self-Rewarding Vision-Language Model via Reasoning Decomposition
-
GPT-Image Excels at Style Transfer Applications
By
–
I disagree with your take on GPT-Image. It's still the best for style transfer.