DeepSeek-V4-Flash maintains four residual streams through its mHC architecture. Each token gets its own read/write weights, so routing should vary with input. But after training, something surprising happens: within the same Attention or FFN sublayer, the dominant stream choice
