I recommend having a list of things AI can almost do well, but still fail at right now These are your personal benchmarks, and the only real way to understand if new models (think GPT-5, Gemini 2.0, etc.) are actually a leap forward in your context or just an incremental change
Personal AI Benchmarks: Measuring Real Progress in New Models
By
–