if you use enough discrete tokens you could def recover performance eventually, unless you’re doin somethin wrong
@jxmnop
-
Multi-token sparsity in discrete token activations
By
–
yeah it’s similar— I think you could call it multi-token sparsity or something like that? I mean obviously the activations have to be sparse in this case but we can use as many discrete tokens as we want (unlike in the projects you’re mentioning)
-
Discrete Tokens vs Continuous: Perplexity Trade-offs
By
–
also in this setup you probably need a much larger number of discrete tokens than continuous ones to get decent ppl
-
Logit Lens Limitations in Early Layer Analysis
By
–
nah that’s not the same thing at all; logit lens is pretty poorly conditioned esp at early layers and certainly doesn’t capture all of the hidden state
-

Latent Text Transformer: Making Model Thoughts Readable
By
–
random research idea: Latent Text Tansformer (LTT) in a nutshell: replace sequence of *vectors* as hidden states of the transformer with sequences of *tokens*, so we can read the model's "thoughts" directly then train a transformer that uses longer sequences of *discrete
-
Google’s Internal Competing Interests and Organizational Structure
By
–
seems more likely that google is just a big organization with lots of groups and competing interests
-

Circuit Diagrams Paper: Delightfully Informative AI Illustrations
By
–
follow him for more delightfully informative illustrations: @vtabbott_ and check out the circuit diagrams paper: http://
openreview.net/pdf?id=RyZB4qX
Egt
… -

Mixtral 7B Architecture Diagram Overview
By
–
bonus diagram of mixtral 7b (from http://
vtabbott.io/mixtral) -

Australian Artist Creates Detailed Neural Circuit Diagrams of Transformers
By
–
today i found out that this one australian guy has been toiling away making incredibly detailed Neural Circuit Diagrams with the vibe of a 1950s issue of Popular Mechanics, but content fit for the 2020s behold. the Transformer
-
Attention Is All You Need: Foundational Paper Reference
By
–
it’s from the original paper – Attention Is All You Need