The AI conversation on X can be frustrating as researchers keep bumping into well-understood problems in economics, sociology, history, & psychology that would be useful to know but are hurt by the lack of dialogue with expets (both as they left X & they aren’t part of AI talk).
@emollick
-

GPT-5 Personality: Sandwich Feedback Better Than GPT-4o
By
–
The new GPT-5 personality likes giving sandwich feedback (you are great- suggestion for improvement – you are great). In general, better than GPT-4o at pushing back while being a bit syncophantic. (It would be good for the AI labs to look at the research on giving good feedback)
-

Benchmarking AI Models for Psychological Safety Risks
By
–
This is a much needed first attempt at a benchmark to measure how much given AI models will play along with users pushing them in delusional or potentially psychologically dangerous directions. Some early signal that full GPT-5 (not chat) is a less psychologically risky model.
-
AI personality becomes key battleground for consumer AI development
By
–
As I predicted (and worried about) AI “personality“ is going to be the battleground for a lot of consumer Ai development. That appears to be the angle so for Grok, and the lesson OpenAI took from the backlash against retiring 4o. It may be consequential.
-
AI Models Leverage Video Features for Viral Adoption Strategy
By
–
It is interesting to see how much effort is going into making ancillary features of the AI models go viral. Ever since the (organic) Studio Ghibli moment, one focus for Grok & Gemini has been on video as a gateway. A challenge has been whether people have creative video ideas.
-

AI Model Development Increasingly Difficult for Highly Capitalized Companies
By
–
Some signs that catching up in the AI model space is rapidly becoming challenging for even the most highly capitalized companies.
-
Pro AI Models Excel at Hard Problems Requiring Expert Evaluation
By
–
The pro models (GPT-5 Pro, Gemini 2.5 Deep Think, Grok 4 Heavy) can be impressive in ways that are hard to see. They take a lot of time to answer questions & are built for very hard problems that require expert evaluation. That is a narrow, but, also very valuable, problem space.
-

GPT-5 Exceeds Medical Professionals on Reasoning Benchmarks
By
–
GPT-4o was below the level of medical professionals on medical reasoning benchmarks GPT-5 (apparently Thinking medium) now far exceeds them. (Usual benchmark caveats apply)
-
Prompt Engineering Impact on Test Outcomes Analysis
By
–
The size of these impacts is much larger than the effect of most prompt engineering on test outcomes.
-

Open Weight Model Performance Varies by Cloud Host Provider
By
–
This is actually a pretty surprising and something that should lead companies to change how they are thinking about hosting. Model performance for the open weights GPT model vary by meaningful amounts depending on who is hosting it, with Azure & AWS being low. Worth watching.