Working on codex at big token helped, learned lots of little tricks to make GPT better and updated the claw harness.
LLMS
-
AI 4.6’s ‘It’s fine the way it is’ response suggests AGI.
By
–
I asked Opus 4.6 to rewrite one of its outputs for me and it responded: "Honestly — it's fine the way it is." Guess we have AGI.
-

Impact of token limits on AI agent performance in cybersecurity
By
–

Real world example: raising token limits from 3M to 10M tripled the amount of work that Codex could do independently on cybersecurity tasks. From 3.1 hours to 10.5 hours.
-

Scaling Laws and Token Usage in AI Reasoning Models
By
–

Unappreciated fact is the second scaling law does not seem to completely plateau in many tasks: throw more tokens at a reasoning AI model and get better answers, especially with a simple harness. Benchmark performance is actually limited by token usage. https://
open.substack.com/pub/joelbkr/p/
many-benchmarks-scores-would-appear?r=i5f7&utm_medium=ios
… -
Policy on Using Claude-P: No Restrictions Needed
By
–
I wouldn't have a policy that says you're not allowed to use "claude -p"
-
Major progress on automated testing with AI assistance
By
–
Not quite at dark factory levels, but that's a big step forward. Brainstormed with codex and it… just works (okay took 6h to build but damn. Way better than old-school e2e tests)
-
New QA Process Using OpenClaw for Self-Testing with Synthetic Messages
By
–
Working on a new QA process where we use openclaw to QA openclaw with a new synthetic message channel. Orchestrator agent understands project, defines tasks (e.g. ask agent to create cron) then verifies that agent did what that. If failure -> spin up subagent to analyze+fix.
-
Gemma 4 E4B On-Device LLM Performance Evaluation
By
–
Gemma 4 E4B is impressive for an on-device LLM. GPT-4ish quality, and expect hallucinations.
— Ethan Mollick (@emollick) 5 avril 2026
Here is: “List five sociological theories starting with u and what they are. Then describe them in a rhyming verse”
Its in real time, the last is a little bit of a stretch, but not bad! pic.twitter.com/nJ5HuQsFSBGemma 4 E4B is impressive for an on-device LLM. GPT-4ish quality, and expect hallucinations. Here is: “List five sociological theories starting with u and what they are. Then describe them in a rhyming verse” Its in real time, the last is a little bit of a stretch, but not bad!
-

Emotion vectors influence Claude AI behavior and task performance
By
–
Going to try to see if instructing Claude to keep calm makes performance on tasks better. Really interesting research from @AnthropicAI on emotion vectors and how they change behaviors of the model. They found that reducing Claude's internal "calm" vector increased blackmail behavior. Boosting "desperation" caused invisible cheating on coding tasks.
→ View original post on X — @jiquanngiam, 2026-04-05 17:57 UTC
