The way "Generative Retrieval" is used these days is to encode documents in a LLM and retrieve it somehow from memory. Technically nothing is ever new according to Schmidhuber (lol)…but as far as this new context and way this term is being used, DSI is the pioneerinng work.
@yitayml
-
Generative Retrieval and Dense Search Index Explained
By
–
lol? Generative retrieval *is* DSI. I don't think you know what generative retrieval means.
-
Experiments Over Writing: The Real Scientific Work
By
–
what? the experiments are the real work though.. writing the paper is the leisure (and fun) part where you can even just do it from a coffee shop with a 13 inch macbook cruising at 10% mental bandwidth.
-
Over-claiming in AI research raises credibility concerns
By
–
Yeah i don't even need to read the paper to know that there is some form of over-claiming involved.
-

LLMs Statistical Bias in URL Generation and Citation Patterns
By
–
It makes sense no? Explanation: this video url probably appears so many times under this kind of context. It's probably one of the most highly "cited" youtube urls… I think most LLMs would produce this particular URL just because of statistics and not because it's trolling
-

Generative Retrieval: Google AI Research Breakthrough Discussion
By
–
Interesting paper from my ex-colleagues at @GoogleAI led by @vqctran
. Generative retrieval (i.e., DSI) is one of the most fun works I've worked on (and pioneered) during my Google career. Also, @vqctran is driving a lot of the agenda that we worked on together back then. He has -
Big-Bench: Novel AI Tasks Beyond Benchmark Aggregation
By
–
Fwiw, Big-bench actually proposed new novel tasks. It's not a repackage. Aggregate results can be broken in many ways as mentioned in the benchmark lottery paper.
-
Meta-Review of LLM Leaderboard Evaluation Methodologies
By
–
In the spirit of being very meta here. Here's my personal meta-review of all the leaderboard-ing methodologies. 1. I like the elo ranking based on chatbot arena from @lmsysorg 2. LM harness (e.g., zero-shot PIQA, Hellaswag etc) is the equivalent of "MNIST" for LLMs. Okay-ish
-
LLM Evaluation Methods: Academic Benchmarks vs Real-World Performance
By
–
Community: Eval for LLMs are broken! Academic benchmarks are not representative of real world performance! . We need better evals! Also the same community: Lets make definitive rankings & leaderboards based on just four zero-shot "LM harness" tasks! Not wanting to single
-
Fine-tune Flan-T5 for your specific use case
By
–
Flan-t5 if you're going to fine-tune for your use case.