I suspect it is all because of the training data and how humans write: the most important information is usually in the beginning or the end (think paper Abstracts and Conclusion sections), and it's then how LLMs parameterize the attention weights during training. 5/5
LLMS
-
RNN-based LLMs and information retention patterns
By
–
This is quite interesting … 1) I would expect that the opposite is true for, e.g., RNN-based LLMs like RWKV (since it's processing information sequentially, it might rather forget early information) 3/5
-
Transformer Architecture and Middle Document Retrieval Performance Bias
By
–
2) To my knowledge, there is no specific inductive bias in transformer-based LLM architectures that explains why the retrieval performance should be worse for text in the middle of the document. 4/5
-

Long Context LLMs: Promise and the Middle Information Problem
By
–
We have seen a new wave of LLMs for longer contexts: 1) RMT, 2) Hyena LLM & 3) LongNet There are several use-cases for such long LLMs but the elephant in the room is: How well do LLMs use these longer contexts? Turns out not so well if info is in the middle of the input.
1/5 -
LLMs Struggle Retrieving Information from Document Middle
By
–
According to "Lost in the Middle: How Language Models Use Long Contexts" (
https://
arxiv.org/abs/2307.03172) LLMs are good at retrieving information at the beginning of documents. They do less well in terms of retrieving information if its contained in the middle of a document. 2/5 -
Lab Advances Robotic Learning with Large Language Models
By
–
Our lab’s new work in robotic learning using #LLMs 👇🤩 https://t.co/aQN130Wozv
— Fei-Fei Li (@drfeifei) 8 juillet 2023Our lab’s new work in robotic learning using #LLMs
-
GPT-4 Remains Most Capable AI Model Despite Competition
By
–
wdym by "they dont present a business threat"? GPT4 is still the most capable model around by a mile.
-
Chatbots as Therapists: Ethical Concerns and Medical Misrepresentation
By
–
i've written b4 about how some people are finding chatbots helpful for therapeutic reasons, but as @dr_ilardi told me, the positioning of the bot as a "psychologist" is worrisome because it refers to a trained medical professional. this, he said, "almost certainly is not that.”
-

16M Chatbots Created: Users Embrace AI Conversation Technology
By
–
people are using it to create chatbots (over 16M so far) and chat with them (some, like a Mario chatbot and a Raiden Shogun and Ei chatbot, have been sent many millions of messages). i spent a lot of time talking to a farting unicorn and albert einstein.
-

When Will LLMs Achieve Mainstream Adoption Beyond Developers?
By
–
OK, I see a ton of LLMs being developed but none have really gotten adoption outside of small group of AI developers. When will that change? @BrianRoemmele we need a guide to help us figure out the pros and cons of each model. Have you seen anything like that?