Not only are there too many papers, but the papers themselves are too much.
@emollick
-
GPT-4 Enterprise Adoption Challenges Across Business Functions
By
–
Just spoke to a biomedical company deploying GPT-4 widely. You can see what makes AI adoption both powerful & hard for firms: people are using it for all sorts of things from marketing to competitive analysis to summarizing scientific papers. You can't just do one-size deployment
-

GPT-5 Hallucination Rates: Continuing the Trend of Larger Models
By
–
One of the big questions about GPT-5 is whether it continues the trend that larger models hallucinate less For example, in this study of medical citations, GPT-3.5 had a hallucination rate of 55%, while GPT-4 (without internet access) made up 18% of cites https://
nature.com/articles/s4159
8-023-41032-5
… -

GPT-4 Lacks Reverse Image Search Capability
By
–
Pretty good, especially as GPT-4 doesn’t have reverse image search & this appears to be a recent meme.
-
Code Interpreter: Balancing Statistical Risks and Accessibility
By
–
As someone who does lots of quantitative work & knows how it can go wrong, I totally get why stats folks are worried about the misuse of Code Interpreter, but would love to see more engagement with how people with less experience can use it effectively. It is an opportunity, too
-

GPT-4 Task Performance Variability: Why Some Configurations Succeed
By
–
This paper is about LLMs solving a specific task, but it is also about the difficulty of figuring out what LLMs do well & why Most configurations of GPT-4 failed to solve the problem, but one robustly did, for reasons that are hard to know. LLMs are weird https://
arxiv.org/pdf/2403.15371
.pdf
… -

GPT-4 Achieves Human-Level Performance in Data Analysis
By
–
Two comparisons of data analysts to Code Interpreter: "Experimental results show that GPT-4 can achieve comparable performance to humans" https://
arxiv.org/pdf/2305.15038
.pdf
… GPT-4 scores over 90% on exams, the data science field is “on the verge of a paradigm shift” https://
arxiv.org/pdf/2307.02792
v2.pdf
… -
Students Using AI: Education Strategies Needed
By
–
What does this mean? I don’t understand why people are interpreting this as “I want to replace writing with AI” – I am saying it is happening among students and we need strategies to deal with it.
-

GPT-4 Vision Medical Scan Analysis: Limitations and Accuracy
By
–
I see a lot of examples of people feeding medical scans into the AI to get results. There is no Claude 3 evaluation I have seen, but tests of GPT-4’s vision capabilities show that it makes a lot of mistakes reading scans. Interestingly, it does quite well on text-based tasks
