That just says it is wrong but doesn’t offer a rationale. I think it is quite a useful way to explain what gtp does and why it is limited. And that is really needed when there is so much hyperbole and over-extrapolation.
LLMS
-
Content-Conditioned Q&A Assistants: A Future Feature
By
–
content-conditioned Q&A assistant is a prominent feature of the future.
-
Ted Chiang analyzes ChatGPT’s underlying mechanisms and implications
By
–
Ted Chiang, a brilliant scifi writer, wrote a brilliant piece on ChatGPT. Well worth reading if you want to understand what’s really going on behind the prompt. https://
newyorker.com/tech/annals-of
-technology/chatgpt-is-a-blurry-jpeg-of-the-web/amp
… -

Step-by-Step Thinking Needs Correct Answers in GPT Models
By
–
One of my favorite results in 2022 was that it's not enough to just think step by step. You must also make sure to get the right answer 😀 https://
sites.google.com/view/automatic
-prompt-engineer
…
(actually a nice insight into a psychology of a GPT; it pays to condition on a high reward) -
Meeting in San Francisco to discuss generative AI and ChatGPT
By
–
I’m in San Francisco for a few days. Hit me up if you’d like to meet and talk about generative AI, cerebral valley, and the profound brilliance & utter stupidity of chatgpt.
-
LLM Collaboration and Chain Rules Innovation
By
–
The thread above was, of course, written in collaboration with one of our powerful Large Language Models. Long live Chain Rules!
-
Token Probability in Language Models Explained
By
–
By smaller events here, we refer to the probability of a token, given past tokens, p(c|ab). In probabilistic language modeling, a “token” is a single unit of text, like a word or part of a word. Modern language models consider a vocabulary size of ~100K tokens.
-
Chain Rule Reduces Token Generation Complexity Exponentially
By
–
To generate a sequence of 1000 tokens requires an insane 100K^1000 = 10^5000 choices. That’s a lot more than the estimated number of atoms in the universe, 10^82! With the chain rule the number of possible choices is "only" 100K * 1000 = 100M, a much more manageable number.
-
Chain Rule of Probability Powers Large Language Models
By
–
The Chain Rule of Probability is a powerful tool behind recent advances in Large Language Models. By multiplying together the probabilities of many smaller events, we can compute the probability of a complex event made up of those smaller events.
p(abc) = p(c|ab) * p(b|a) * p(a) -

Ghost Cites and ChatGPT: Academic Citation Crisis Worsens
By
–
This is a classic story about an article that was never written yet became widely cited. "Ghost cites" were always problem in academia. But it's about to get much, much worse. This is how Chat GPT summarizes this famously non-existent article: