Can a model compose parametric knowledge within a single forward pass without explicit CoT? Well, vanilla transformers often fail. But this paper shows that a recurrent-depth transformer may fix that. It reuses the same transformer block for multiple iterations, instead of
Recurrent-Depth Transformers Enable Parametric Knowledge in Single Pass
By
–
