We partnered with @FireworksAI_HQ to fine-tune an @Alibaba_Qwen judge model to detect “Perceived Error” from user interactions. See our findings
LLMS
-

Fine-tuned models match frontier performance with lower cost
By
–
Fine-tuned models match frontier performance In our research with @FireworksAI_HQ
, a fine-tuned @Alibaba_Qwen outperformed all model sizes. They’re also cheaper to run at scale 10-100x depending on trace volume and model choice As trace volumes grow, so will cost-savings on -
GLM 5.2 MoE, NVFP4 467GB, DGX Station memory, offloading works
By
–
GLM 5.2 is an MoE, NVFP4 is 467 GB, and the DGX Station comes with 496GB LPDDR5X + 252GB HBM3e GPU memory With the right offloading formula, it should work
-
Working on getting GLM 5.2 NVFP4 by Luke Alonso running
By
–
Currently working on getting GLM 5.2 NVFP4 by Luke Alonso up and running 🙂 Will report back
-
Sazabi: AI Agent that Detects and Fixes Production Bugs
By
–
AN AGENT THAT FIXES YOUR BUGS ON ITS OWN
— Nico (@nicos_ai) 25 juin 2026
this is Sazabi and it sounds like science fiction
→ watches your software in production
→ detects what breaks
→ investigates the cause
→ and fixes it with Claude Code or Codex
goodbye 3 AM calls, hello shipping features https://t.co/LdVVDF91GPAN AGENT THAT FIXES YOUR BUGS ON ITS OWN this is Sazabi and it sounds like science fiction
→ watches your software in production
→ detects what breaks
→ investigates the cause
→ and fixes it with Claude Code or Codex goodbye 3 AM calls, hello shipping features -

No solution to AI hallucinations despite years of promises
By
–
Crazy how many times people have told me over the last five years that a solution to hallucinations was right around the corner — and yet here we still are.
-
Why Claude Code outperforms general LLM apps with grep and search
By
–
It just greps a lot and searches that's why Claude Code is better at it then a general LLM app
-
Using AI to psychoanalyze yourself from years of chat logs
By
–
A fun exercise is exporting your 5+ year of chat logs and putting it in Claude Code to do psychoanalysis of yourself
-
User finds local LLMs painfully slow on single RTX 5090
By
–
And yes I've tried local LLMs but with just 1x RTX 5090 it's painfully slow and useless
