Re-passing activations through the same layers sounds like recurrent computation stitched back in, curious if it still benefits from longer contexts the way dense transformers do.
Recurrent Computation in Transformers: Long Context Benefits Analysis
By
–