pretty happy with my current coding setup: – gpt-5.2-codex
– xhigh for planning
– medium for implementation
– background terminal (parallel jobs)
– /fork conversations to explore multiple variations of a problem
– codex web via slack for quick async validation/ simple changes
–
LLMS
-
Practical AI-Driven Coding Workflow and Tooling Setup
By
–
-
New video: complete MCP and AI Agents setup guide
By
–
MCP & AI Agents 101: full setup guide New YT video dropped
-

Lang-ladies SF Breakfast: Women Building with LangChain Tools
By
–
Lang-ladies SF Breakfast Join us January 30th for breakfast + coffee with women building with LangSmith, LangChain and LangGraph. Connect, collaborate, and talk agents, LLM apps, & real-world AI. LangChain office in SOMA 8-9:30am We have a couple of seats left— RSVP
-
Using the Most Expensive Model Available
By
–
Whatt do you mean? we literally use the MOST EXPENSIVE model
-

Open-source Kimi 2.5 on own infrastructure with no limits
By
–
Damn, but open-source models like Kimi 2.5 on your own infrastructure WITH NO LIMITS AT ALL, but WHAT A BLAST !!!!! That's some crazy shit !!!!!
-
Gemini’s Remarkable Growth Rate and Future Trajectory
By
–
The growth rate of Gemini is truly remarkable. If you model the current growth of the different alternatives and extrapolate into the future, it's very clear where this is going.
-

CooperBench: AI Agents Perform 50% Worse in Teams
By
–

Introducing the curse of coordination. Agents perform 50% worse in teams than working alone. People building human-AI collaboration today don't realize why current LLMs fail to be good teammates. We built CooperBench to study this. For humans, we recognize that teamwork isn't just the sum of individual capability. Communication and coordination often outweigh raw skill. But for AI? We're only hill-climbing benchmarks that evaluate solo technical abilities. CooperBench A benchmark to evaluate agent cooperation in realistic software teamwork tasks. The setup is intuitive: two agents, two tasks, two VMs, one chat channel (agents can send over arbitrary text, even the entire patch they wrote). We evaluate whether the merged solution from both agents passes the requirements of both tasks. The curse of coordination The most striking result: agents perform 50% worse in teams (black line) than working alone (blue line). Why is this happening? Is it because they can't use the communication tool? No. They spent 20% of their time sending messages. The problem? Those messages were repetitive, vague, ignored questions, or straight-up hallucinated. But bad communication is only part of the story. We found two deeper failures: Commitment: Agents don't do what they promised. Expectations: Agents don't expect others to keep promises either. Without these, cooperation collapses. However, there is a silver lining We also find emergent coordination behaviors, e.g. role division, resource division, and negotiation, which gives us hope that we can use reinforcement learning to improve coordination. What's next? It is true that highly-engineered multi-agent orchestration could largely sidestep the coordination problem. However, we care more about the AI's capability: if we truly want AI to be our teammates, we need them to be natively capable of effective communicating and coordinating. Two agents on software tasks is just the beginning. The real goal: agents that can cooperate with us well enough to actually empower us. CooperBench is our first step. If you're working on this too, let's talk.
-
Comparing AI Detection Accuracy: GPTZero vs Bypass AI
By
–
Comparez les résultats de ses tests avec ceux de GPTZero :
— Jouhatsu | AI Influence Operator (@Jouhatsu_ai) 28 janvier 2026
Texte de ChatGPT : 100% AI Generated
Texte Bypass AI : 3% généré par l'IA pic.twitter.com/7RNCoGNyjXComparez les résultats de ses tests avec ceux de GPTZero : Texte de ChatGPT : 100% AI Generated
Texte Bypass AI : 3% généré par l'IA
