models will undoubtedly get to that point, and the METR benchmarks will undoubtedly shift to a frame above their current one—of which there are many
@danshipper
-
Critiquing the practice of testing AI tools for failure
By
–
“We got a tool to perform poorly” is the lowest form of science and journalism imo and is only relevant when the tool is, in fact, extremely useful
-

How Human Prompting Influences AI Model Benchmarks
By
–
mythos obviously looks incredibly capable and im psyched to use it also if you're panicking about it: benchmarks don't measure model capability alone they measure model capability after a human has done the work of finding a prompt that lets the model’s capability appear that
-
The rising value of human-AI creative collaboration
By
–
as ai makes imitation cheaper and cheaper the value of using AI and your brain to make totally new things goes up
-
Managing multiple AI coding threads and workflows
By
–
to be clear, im not even running inference. just like, 5 Codex threads + 1 Claude Code thread plus a few other things
-
Running local AI agents reveals hardware compute bottlenecks
By
–
running agents on my laptop is the first time in a long time where i feel my computer is underpowered i could actually consume way more RAM and GPU if it was available
-
Popular AI agent frameworks for development
By
–
Custom (just a python file) and Viktor seem to be most popular right now
-

Using AI tools to generate a pre-game podcast
By
–
having codex + claude make a pre-game podcast for me before my writing session today lol
-

Using AI for music transcription
By
–
happy saturday! using codex to analyze and transcribe the piano arrangement of a youtube music video i love
-
Investment opportunities in the AI market
By
–
Generational opportunity for anyone in AI to play the markets given this time lag. Guarantee everyone is psyched about Codex in a few months. Invest accordingly