I agree that it is a problem that the models have no idea of their own limits, it is one of many issues that make LLMs hard to use. And yes, agree image comprehension and image creation are both limited, but the evidence suggests pretty rapid improvement & some real utility.
LLMS
-

Image Generation Progress: Spaghetti Forks and LLM Limitations
By
–
Well, I got six forks made of spaghetti on the first try, but one is a double-sided fork It is pretty amazing how far imagegen has come in the past years (they aren't flawless, but this would have been impossible months ago). Yet they aren't really a good measure of LLM ability
-

Why LLMs Should Weird You Out: Understanding Their Capabilities
By
–
I don't know anyone who uses LLMs who is not occasionally weirded out by what they can do. If you are not, you should be. They are weird. Wolfram had a rather startling (at least at the time) theory after using ChatGPT. Understanding whether he is right is important.
-
AI Model Achieves 90% Accuracy on Prompt Tasks
By
–
I have tried both. This one just makes the flow from a prompt and get things right 90% of the time.
-
The Deep Mystery: How LLMs Simulate Human Thought
By
–
We really have not made a lot of progress on explaining the deep mystery of LLMs: How does a model using matrix multiplication to predict the next word manage to simulate human thought well enough to do all the very human-like things it does? And what does that mean about us?
-
How Do LLMs Simulate Human Thought Despite Small File Size
By
–
I think it is actually makes LLMs even weirder! The next question is "how does a file the size of a moderately sized video game simulate human thought" and I don't think we have good answers.
-
MiniCPM-V 4.5 Chat App Development with Impressive Benchmark Performance
By
–
vibe coding a MiniCPM-V 4.5 @OpenBMB chat app in anycoder
— AK (@_akhaliq) 28 août 2025
MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro,… pic.twitter.com/r1i5b6JfFpvibe coding a MiniCPM-V 4.5 @OpenBMB chat app in anycoder MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro,
-
LLMs Demystified: Matrix Multiplication and Local Inference
By
–
I use this line about LLMs a lot – that they're a bunch of matrix multiplication – because I think it demystifies them A multiple GB file of floating point numbers you can download and run matrix multiplications against on your own machine is a bit less weird and frightening
-

Self-Rewarding Vision-Language Model Through Reasoning Decomposition
By
–
Self-Rewarding Vision-Language Model via Reasoning Decomposition
-
LLM Fake Coding The Dark Side of AI Employment
By
–
The killer app for LLMs is Fake Coding. It’s all you need to land multiple simultaneous high-paying jobs while going to the beach. And when they finally fire you, you (i.e., your bot) can use your now-golden resume to land your next batch.
