ViT was designed by Dosovitskiy, @giffmana
, and other researchers from Google Brain. The discussed paper can be found here:
MULTIMODAL AI
-
ViT Vision Transformer designed by Dosovitskiy and Google Brain researchers
By
–
-

Vision Transformer Paper Explained – New Video Tutorial
By
–
Vision Transformer(ViT) Paper Explained – A NEW VIDEO ViT paper introduced transformers for large-scale image recognition. It is one of the influential papers in visual representation learning and modern computer vision in general. In my new video, we discuss ViT paper:
-

Foundation Models for Embodied AI and Robotics Performance
By
–
Is there a foundation model for embodied AI / robotics? Turns out that while today's visual 'foundation models' outperform learning from scratch baselines, there is no FM that is universally dominant across 17 different tasks spanning locomotion, navigation, dexterous, and
-
Machine Learning Detects Material Similarities for Robotic Scene Recognition
By
–
Scientists are using machine learning to detect similar materials in images, opening up new possibilities in robotic scene recognition & research. By training algorithms on large datasets, researchers can identify patterns & similarities. Read more: https://
bit.ly/3Clr59z -
Stable Diffusion Fine-Tuned to Midjourney Aesthetic Levels
By
–
We've now fine tuned a base stable diffusion model to Midjourney levels of aesthetics (particularly in some domains).
-
ARTIC3D: Learning Robust Articulated 3D Shapes from Web Images
By
–
ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections
— AK (@_akhaliq) 8 juin 2023
paper page: https://t.co/UGFpZ1Rce6
Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture,… pic.twitter.com/nbhred7HDPARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections paper page: https://
huggingface.co/papers/2306.04
619
… Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, -

Designing a Better Asymmetric VQGAN for Stable Diffusion
By
–
Designing a Better Asymmetric VQGAN for StableDiffusion paper page: https://
huggingface.co/papers/2306.04
632
… StableDiffusion is a revolutionary text-to-image generator that is causing a stir in the world of image generation and editing. Unlike traditional methods that learn a diffusion model in -

Youku-mPLUG: 10M Chinese Video-Language Dataset for Pre-training
By
–
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks paper page: https://
huggingface.co/papers/2306.04
362
… To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we -

Text-only Domain Adaptation for Speech Recognition using Unified Representation
By
–
Text-only Domain Adaptation using Unified Speech-Text Representation in Transducer paper page: https://
huggingface.co/papers/2306.04
076
… Domain adaptation using text-only corpus is challenging in end-to-end(E2E) speech recognition. Adaptation by synthesizing audio from text through TTS is -

M3IT: Large-Scale Multi-Modal Multilingual Instruction Tuning Dataset
By
–
M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning paper page: https://
huggingface.co/papers/2306.04
387
… Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks.