Interesting, it somewhat confirms my intuition: > "even with a longer time horizon, xhigh doesn't solve significantly more tasks" Models often hit a conceptual wall and in those cases no amount of extra time will help!
LLMS
-
Model Inference Performance and Tool Call Optimization Analysis
By
–
OK, I will ponder it! Currently: it's only one turn, about 20-30 tool calls (est.) depending whether you include file reads, and networking is definitely not the bottleneck. But yeah, load/inference speed is punished — but that's real-life! I think Kimi K2.5 might have
-
Codex 5.3 Release Generates Strong Momentum in Developer Community
By
–
codex momentum is strong, and many people are feeling just how big of a leap 5.3 is. if your organization hasn't tried codex yet, it's worth revisiting.
-

Minimax M2.5 Launch: Advanced AI for Productivity Workflows
By
–
Minimax M2.5 is live on Poe! The new Minimax M2.5 is designed for real world productivity, with a focus on planning‑driven, multi‑step execution across complex digital workflows. It excels at coding, cross‑tool context switching, and agent style task coordination. Also, it
-
Peter joins OpenAI: Strategic competition between OpenAI and Anthropic intensifies
By
–
这次 Peter 加入 OpenAI,热闹之外,几个值得思考的点: 1. 从最开始的 Clawdbot 到 Moltbot 再到 OpenClaw,名字就可以看出来,Peter 开始是 A 社模型的粉丝 ,到现在去 OpenAI 背后的思考。 2. OpenAI 和 A 社之间,现在不只是模型在打,社媒、关注度、开源也在打。可以理解为,OpenAI
-
Speculation on Model Release Strategy and Feature Deployment
By
–
The clawdbot —> open claw rename foreshadowed it all. Zuck must not be too happy. And interesting that Anthropic didn’t even make a play. So what does this mean? I suspect new functionality keeps coming to open claw first – and the best stuff graduates to chatgpt proper. A
-

Open-source multimodal RAG framework discovery and integration
By
–
Ever since I started working on Memory, I've been seeing RAG products every day. Discovered a one-stop RAG framework, open-source, MIT licensed. This library can be considered a multimodal superset based on LightRAG. Building on the LightRAG architecture, it provides a
-

Open Weights Models Struggle on Logic Reasoning Benchmark
By
–
@nrehiew_ Hey, just made a logic reasoning / problem solving benchmark where open weights models get completely lost, but the frontier models make it look easy. Curious about your hypothesis why, thinking it's sparsity related:
-

Open Models Overfitting Benchmarks While Losing Reasoning Ability
By
–
@xeophon On the topic of swe-rebench and lower scores, another data point for you: my own analysis suggests open models are overfitting to popular patterns/benchmarks while failing to get better at logical reasoning / problem solving:
-

Logic Reasoning Benchmark: Frontier vs Open Weight Models
By
–
@scaling01 Before your pivot to Star Wars memes, I remember you used to be interested in LLMs! I just built a logic reasoning / problem solving benchmark where frontier models one-shot solutions, but the open weights models really struggle: