I agree that it is a problem that the models have no idea of their own limits, it is one of many issues that make LLMs hard to use. And yes, agree image comprehension and image creation are both limited, but the evidence suggests pretty rapid improvement & some real utility.
GENERATIVE AI
-
Nano Banana AI Model Disrupts Photo Editing Market Leadership
By
–
This is the first AI model that really disrupts the photo editing market.
— Allie K. Miller (@alliekmiller) 28 août 2025
I don't think Nano Banana replaces Photoshop, but it sure as heck replaces something.
My immediate guess: Facetune is about to lose users by the thousands. It maintains people and objects really well. pic.twitter.com/k47kKhYp8IThis is the first AI model that really disrupts the photo editing market. I don't think Nano Banana replaces Photoshop, but it sure as heck replaces something. My immediate guess: Facetune is about to lose users by the thousands. It maintains people and objects really well.
-
AI Vision Models: Weaknesses in Counting and Image Generation
By
–
Clear weak spots remain counting, generating alternate images when the training data is thick (full glasses of wine, clocks with oddly set hands), etc. It isn't hard to make them fail. But there is a lot they do very well, and the gains have been pretty quick so far.
-

Image Generation Progress: Spaghetti Forks and LLM Limitations
By
–
Well, I got six forks made of spaghetti on the first try, but one is a double-sided fork It is pretty amazing how far imagegen has come in the past years (they aren't flawless, but this would have been impossible months ago). Yet they aren't really a good measure of LLM ability
-

Why LLMs Should Weird You Out: Understanding Their Capabilities
By
–
I don't know anyone who uses LLMs who is not occasionally weirded out by what they can do. If you are not, you should be. They are weird. Wolfram had a rather startling (at least at the time) theory after using ChatGPT. Understanding whether he is right is important.
-
AI Model Achieves 90% Accuracy on Prompt Tasks
By
–
I have tried both. This one just makes the flow from a prompt and get things right 90% of the time.
-
How Do LLMs Simulate Human Thought Despite Small File Size
By
–
I think it is actually makes LLMs even weirder! The next question is "how does a file the size of a moderately sized video game simulate human thought" and I don't think we have good answers.
-
MiniCPM-V 4.5 Chat App Development with Impressive Benchmark Performance
By
–
vibe coding a MiniCPM-V 4.5 @OpenBMB chat app in anycoder
— AK (@_akhaliq) 28 août 2025
MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro,… pic.twitter.com/r1i5b6JfFpvibe coding a MiniCPM-V 4.5 @OpenBMB chat app in anycoder MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro,
-
AI System Achieves Perfect Performance in Single Attempt
By
–
I agree. I was super impressed with how it got everything right in just one shot. Haven't seen anything like that yet.
-
Get 400 Free Credits for Lindy AI Tool
By
–
Try out @getlindy now and get 400 credits absolutely free:
