i think i see your point. the embeddings we work with are the average of many tokens’ last-layer activations. so we can’t directly load them back in, need to train a model to do it
Training Models to Reconstruct Token Embeddings from Activations
By
–