their neural networks worked better than any other methods and the bigger ones are better although at this scale, mixing with ngrams still helps a lot. they use a "mixture of models" – similar to today's MoEs, but the experts are different ngram models, plus one neural network
Neural Networks Outperform N-gram Models in Mixture Architecture
By
–
