Researchers have just explained why massive neural networks work so effectively. Deep networks possess enough capacity to memorize completely random noise. Yet, they still excel at making accurate predictions on unseen data. A new paper finally sheds light on this phenomenon. The empirical Neural Tangent Kernel (NTK) subtly divides the output.
