Eventually tracked down the problem to the Unsloth monkey-patching. Distributed runs even with one process fail, but if you use the regular models from transformers and `pef`, trained with `trl` then it doesn't deadlock… Hmm.
Unsloth Monkey-Patching Causes Distributed Training Deadlock Issues
By
–