

Huge! We are finally moving past the era of using separate vision encoders and generative models to build multimodal AI! SenseTime and NTU just introduced NEO-unify,a native, unified, end-to-end paradigm. Instead of using middleman tools to translate images, this model https://
x.com/SenseTime_AI/s
/SenseTime_AI/status/2029585218819199108
…
