Claude Code Chrome Extension FTW!
AI
-
The irresistible urge to announce AI labs train on my benchmark
By
–
The thing I most need a name for right now is that irresistible urge to announce to myself that AI labs are probably training on my benchmark as soon as I share an image of a pelican riding a bicycle anywhere on the internet.
-
Shader test as a measure of AI coding capability
By
–
Except a shader like this is a very good measure of model capability because of the technical difficulty of building this sort of code. It translates to other coding as well. Feel free to see my many other tweets and substack posts (& book) about AI applications in businesses.
-

AI Agent Autonomously Automates PR Merge Workflow
By
–
While working on openai/codex today, Codex surprised me by setting up its own heartbeat automation to poll my PR until CI was green and approvals were in. It kept checking every 10 minutes for an hour, merged the PR once everything was ready, then removed the automation. High
-
Edge-Cloud Hybrid Architecture for AI Development
By
–
The architecture is typically hybrid: edge handles latency-sensitive control, cloud platforms handle analytics and AI development at scale.
-

Client Spent $500M in a Month After Missing Claude Usage Limits
By
–
NEW: AI consultant reveals a client accidentally spent $500,000,000.00 in a single month after failing to set employee limits on Claude usage.
-

AI Wins Gold at Math Olympiad via Simple Unified Scaling
By
–
Cool! AI can win gold at the International Math Olympiad via Simple and Unified Scaling! Researchers from Shanghai AI Lab, CUHK, Tsinghua and PKU introduce SU-01. Their simple recipe: first train on proof-search and self-checking behaviors, then scale via two-stage
-
Half of reactions would be LLM-related psychosis
By
–
I think that half of them vaguely correspond to a psychosis linked to LLMs!
-

AI Reasoning Laziness and Prompt Injection Solutions
By
–
Anthropic semble avoir trouvé une solution aux problèmes où l’IA vous donne une réponse incorrect car elle a la flemme et préfère arrêter le raisonnement. Moi je dis ils doivent faire du prompt injections à chaque thinking avec des menaces .
-

Anthropic Acceleration: Model Release Cycle Shortening Trend
By
–

Anthropic released the next version sooner than I thought – the trend is accelerating – from 50-70 days before, down to 42 days since Opus 4.7