One of the starkest examples of how technological innovation changes the kinds of jobs available (and how pre-Industrial Revolution life was very different).
@emollick
-

O3 Model Tests Scientific Paper Error Detection at 21% Accuracy
By
–
What happens if you put a full scientific paper into AI and ask it to find known errors in proofs, tables, etc? Every model before o3 fails completely, o3 gets 21% (its better at proofs, worse at tables & figures). Progress & perhaps a second opinion, not yet autonomous science.
-

AI Productivity Gains Proven Real Users Survey Data
By
–
The repeated argument that AI is not actually useful to real people needs to be retired based on the representative national surveys we now have on real AI users. Teachers using AI report 6 hour a week time savings. Workers using AI report 3x productivity gains on 1/5 of tasks.
-
Building AI Solutions for Tomorrow’s Cost-Performance Curve
By
–
Many firms built around the limitations & cost assumptions of GPT-3.5 class models, and are now stuck with complex solutions that are more expensive & worse than a reasoner without any scaffolding You need to build solutions with an eye towards riding the cost/performance curve.
-

AI Agents Brand Preferences and Advertising Influence
By
–
AI agents have brand “preferences” and are attracted to different kinds of ads (Operator is a fan of buying whatever Bing advertises, Claude has different preferences) There is likely going to be a lot of money spent trying to influence this in the nearish future.
-
Making LLMs Work: Beyond Out-of-the-Box Limitations
By
–
Studies of LLMs keep looking at the (very real) failures of LLMs working out-of-the-box in complex use cases. I am always surprised that naked LLMs can get so far as generalist systems But if you want to really make an agent handle a complex workflow, you can often figure it out
-
Practical Solutions for AI Agent Reliability and Error Reduction
By
–
In practice, for many useful applications, many of the various obvious problems with AI agents (drift, hallucination, compounding errors) are more solvable than they are in theory Clever prompting, tool use, constrained topics,
LLM judges & organizational process close some gaps -
Non-technical experts excel at AI prompt engineering and product development
By
–
Many of the best prompters I have met who are creating actual useful products in organizations are not technical. In fact, coders often struggle with non-deterministic systems in a way that teachers and managers do not. Broaden AI development beyond engineering.
-
Two Paths to Mastering AI: LLM Understanding and Instruction Design
By
–
All the technical language around AI obscures the fact that there are two paths to being good with AI:
1) Deeply understanding LLMs
2) Deeply understanding how you give people instructions & information they can act on. LLMs aren’t people but they operate enough like it to work -
AI UX design needs embrace variance and branching exploration
By
–
When there is a lot of natural randomness and discovery in an AI use case (image creation, innovation), the focus should not be on single-threaded conversation that becomes self-reinforcing through autoregression, but embracing variance, randomness & branching. Calls for new UX
