3/ V-JEPA – vision models trained on a feature prediction objective using 2 million videos; relies on self-supervised learning and doesn’t use pretrained image encoders, text, negative examples, reconstruction, or other supervision sources.
V-JEPA: Self-Supervised Vision Model Trained on Two Million Videos
By
–
