One thing I've never understood is how much value an individual document has to training a model – are there single books that, if included in the training data, would
materially improve the resulting model?
Document value in model training data
By
–