DAM learns scaling coefficients for each column in the weight matrices of the source models. The objective function is optimized with gradient descent. It includes: – KL divergence to align the merged model's predictions with each expert source model
– Cosine similarity
DAM: Scaling Coefficients for Source Model Weight Merging
By
–
