You know how AI benchmarks seem to measure intelligence on abstract problems, but then it turns out the models can't even think logically? This one instead focuses on problem solving creativity, not nerdy math… NEW: The Joy Of Benchmarks Q1'26 is out!
@alexjc
-
Zai API reliability concerns following GLM5 launch
By
–
This is a benchmark-specific timeout, I terminate them if they get stuck with no sign of making progress. The Zai API has not been so reliable since GLM5 launch, but much better in the past 24h-36h.
-
Cursor AI struggles with tool reliability and context management
By
–
Not very polished… In Cursor, it barely worked at all even! Couldn't even call basic tools reliably, started screwing up files at turn two as if it was at the end of its context.
-
Compute Budget vs Token Count: Scaling Model Performance
By
–
Yeah, I get it. Just like pass-k also improves things a lot when you increase k, predictably so! Thinking a compute budget is the best compromise in this case, as it's a bit more grounded & less biased than token counts…
-
Models Hit Conceptual Walls Beyond Extended Reasoning Time
By
–
Interesting, it somewhat confirms my intuition: > "even with a longer time horizon, xhigh doesn't solve significantly more tasks" Models often hit a conceptual wall and in those cases no amount of extra time will help!
-
Model Inference Performance and Tool Call Optimization Analysis
By
–
OK, I will ponder it! Currently: it's only one turn, about 20-30 tool calls (est.) depending whether you include file reads, and networking is definitely not the bottleneck. But yeah, load/inference speed is punished — but that's real-life! I think Kimi K2.5 might have
-
Continual Learning Benchmark: Knowledge Transfer for AI Models
By
–
Thanks, still finishing the blog post so I'll cover all that! It's designed closer to a continual learning benchmark, new context for each problem and models get to transfer their lessons to future selves. They also have access to prior solutions, as that handoff is necessary to
-

Open Weights Models Struggle on Logic Reasoning Benchmark
By
–
@nrehiew_ Hey, just made a logic reasoning / problem solving benchmark where open weights models get completely lost, but the frontier models make it look easy. Curious about your hypothesis why, thinking it's sparsity related:
-

Open Models Overfitting Benchmarks While Losing Reasoning Ability
By
–
@xeophon On the topic of swe-rebench and lower scores, another data point for you: my own analysis suggests open models are overfitting to popular patterns/benchmarks while failing to get better at logical reasoning / problem solving:
-

Logic Reasoning Benchmark: Frontier vs Open Weight Models
By
–
@scaling01 Before your pivot to Star Wars memes, I remember you used to be interested in LLMs! I just built a logic reasoning / problem solving benchmark where frontier models one-shot solutions, but the open weights models really struggle:
