Baffled by the number of people (specifically influencers and investors) who still get fooled by ML demos and cherry picked examples engineered for wow effect! Rookie mistake to judge the quality of a model (versus looking at usage, revenue and more tangible evaluations)
AI
-
GPT-3 cannot perform accurate calculations without step-by-step writing.
By
–
In general, you shouldn't expect it to be able to perform accurate calculations, at least not "in its head" — this is a known limitation of models like GPT-3. It only stands a chance if prompted to write calculations out step-by-step like one might do on paper.
-

ChatGPT avoids hallucination on Hofstadter/Bender questions
By
–
You're talking about a different model — this post is about ChatGPT. text‑davinci‑003 still fails on all of the Hofstadter/Bender questions that I've tried. The prompt you're suggesting does not seem to produce a hallucination in ChatGPT:
-
AI Will Handle Boring Tasks in Future Jobs
By
–
In the future, all the boring parts of your job will be done by AI.
-
Turing test and human false confidence about cars in space
By
–
Passing the Turing test isn’t really the goal — e.g. it doesn’t try to fake a realistic amount of ignorance about the birthdates of US presidents. But if it were, it would pass here: many humans would tell you, confidently, that there are no cars in space.
-
Challenge of incorporating all known trivia into training data
By
–
Expecting it to fully integrate every piece of trivia seems unfair — a lot of people would rate the first answer as reasonable. It’s fundamentally difficult to get training data that incorporates everything the model knows, rather than what the human demonstrator knows.
-
GPT-3 improvement shows not just pattern-matching, still hallucinates details
By
–
The point isn’t that it’s perfect, just that it’s not narrowly pattern-matching on this specific list of questions — there’s clear improvement vs. GPT-3 across many questions that contain false assumptions. It does still hallucinate details in other ways.
-
NeurIPS Rejects Sound Papers on Ethics Grounds
By
–
NeurIPS continues to reject scientifically sound papers on “ethics” grounds. How much longer will we put up with this?
-
Tokenization issue: model struggles with letter sequences
By
–
This feels like a tokenization issue — in general, it doesn’t see text as sequences of letters, and struggles with tasks that assume it does.
-
NeurIPS hybrid conference format criticism virtual talks live posters
By
–
Apparently some people’s idea of a good hybrid conference is that all the talks are virtual and only the posters are live. Not kidding: that’s NeurIPS this year.