I don't think you need databases to pretrain LLMs because you can basically throw out the embedding database after each batch since the embedding layer changes after each training iteration.
No Need for Databases in LLM Pretraining Due to Embedding Changes
By
–