Did you try any different configurations of the flow model than 4 layers? I would generally expect a wider 2 layer to train faster, unless there is some character to the flow problem that needs more abstraction.
Flow Model Architecture: Exploring Layer Configurations and Training Efficiency
By
–