Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate (first screenshot, Kalra and Barkeshli): https://
arxiv.org/abs/2605.21486 Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size (Hayou and Liu): https://
arxiv.org/abs/2506.15025
Quantifying Hyperparameter Transfer and Embedding Learning Rate in LLMs
By
–