DeepSeek continues to have the most unhinged reasoning traces (“pacing mentally”)
@emollick
-

AI Systems Collapse Traditional Writing-Based Proof Systems
By
–
As I wrote a couple years ago, systems assuming writing can act as a proof of effort, ability, or care are going to collapse and need to be reconstituted. Everything from letters of recommendation to expert reports to essays to performance reviews… https://
oneusefulthing.org/p/setting-time
-on-fire-and-the-temptation
… -
Model Updates, Prompt Testing, Edge Cases and AI Security
By
–
Other questions: When do you update models? How are you testing your prompts? Have you tested for edge cases and biases? Can we vet the prompts you use? What happens if an AI service goes down? How are you working within context windows? How are you dealing with prompt injection?
-
xAI transparency and RAG limitations need greater disclosure
By
–
I am a broken record on this, but if truth-seeking is really an xAI value, they need to be much more transparent about what their system can do and the limits of their RAG approach, especially as implemented in X. (Also, system card!)
-
AI Delegation Risks and the Role of Formal Education
By
–
I think there is reason to be concerned that AI allows full-scale delegation of thinking and writing tasks that are much bigger than memorizing numbers. The answer is likely to be formal education – places where you can make people go through the hard mental work of learning.
-

Technology Outsources Our Cognitive Processing Abilities
By
–
Every new form of information technology makes us dumber in very specific ways, as we outsource some of our processing to the technology. People seemed to think cellphones might make us dumber as we didn’t have to remember numbers (but that remembering passwords would save us)
-
Gemini’s Video Processing: Multimodal AI for Content Analysis
By
–
Gemini is good at processing video (using frequent screenshots & audio transcripts). I gave Gemini a video on a historical recipe, it was able to find visual elements not mentioned in the transcript. It is not hallucination-free, but there are lots of new use cases for screening
-
Grok 3.5 Should Include Model Card Documentation
By
–
Grok 3.5 really needs to come with a model card.
-
O*NET Economic Studies Match AI Adoption Patterns at Anthropic
By
–
Going to disagree with you here. The economic studies using O*NET to predict job overlap are matching well with the actual adoption patterns Anthropic have been reporting.
-
Apple AI Models Perform at Chance Level on GPQA Benchmark
By
–
Isn’t GPQA a 4 option multiple choice test? Apple does no better than chance (as do other small models)
