I don't actually like classic "unit tests" – my preference is for the kind of tests people often call "integration tests" that show feature works end-to-end I usually try to say "automated tests" to avoid that unit test confusion
CODE
-
Understanding Agent Concepts with ADK
By
–
It's good content to understand the Agent concepts with ADK.
-
Using ChatGPT to analyze podcast transcripts and AI content
By
–
Something I am experimenting with. I copy pasted: 1) the full podcast transcript
2) the bitter lesson blog post
3) my full post above To ChatGPT. The interesting part is you can fork the conversation context to ask any questions and take it in whatever direction with chat: -
Documentation and Testing: Essential Software Development Practices
By
–
Without tests you can't be sure your features still work after you make changes to them Without documentation you can't effectively test the software because you never wrote down what it was meant to do in the first place!
-

DeepEval: Open Source LLM-as-Judge Evaluation Framework
By
–
Evaluate multi-turn conversations with just a few lines of code! DeepEval lets you build decision-tree based LLM-as-a-judge evals that break down complex chats step by step. 100% Open Source.
-
Claude 4.5 vs ChatGPT 5: Best Use Cases
By
–
After running all these tests, here’s my advice: – For apps, websites, and complex coding projects → use Claude 4.5
– For general tasks, reasoning, and everyday use → use ChatGPT 5 Both are powerful, but each shines in different lanes. Which one are you going to be using? -
Perfect Commits: Best Practices in Software Development
By
–
I genuinely think they are – I wrote a bunch more about these a few years ago in https://
simonwillison.net/2022/Oct/29/th
e-perfect-commit/
… and then in this talk -
Engineering Practices That Boost Coding Agent Productivity
By
–
It's not just unit tests – there are so many other top tier software engineering practices that accelerate productivity with coding agents Automated tests, comprehensive documentation, good version control habits, a culture of code review, quick deploy to staging environments…
-
AI Code Generation Fails to Improve Shipping Productivity
By
–
The issue is that it's not increasing productivity on *shipping* by 2x — for most people I know I can see it's actually reducing productivity on that metric. The AI code doesn't fit within established software engineering practices and doesn't allow for an effective e2e process
-
EXL3 Quantization Development for AI Models
By
–
working on an exl3 quantization for it right now i am more excited for this though (quote tweet)