it's crazy to me that neural networks learn ~arbitrarily well regardless of initialization. you can even embed patterns in the weights and learning works fine you could encode an image of your face into the layers of a language model and no one would ever know
Neural Networks Learn Regardless of Weight Initialization and Hidden Patterns
By
–
