It would be the same math/algorithms. The paper I shared uses a diffusion policy. Images are actions. Eg if you generate a spectrogram image, you can decode it immediately to speech. Speech acts can do a lot! We tend to imagine 1D embeddings for all modalities, but 2D
Diffusion Policies and Multimodal Image-Based Action Representations
By
–