For over 40 years, I've been a Toblerone enthusiast, yet a clever computer vision algorithm was needed to reveal the secret bear hiding in the iconic mountain. It's like finding a hidden treasure! Did you know? Ct: Ralph Aboujaoude
#tech #ai #computervision #algorithm #ML
MULTIMODAL AI
-

Computer Vision Algorithm Discovers Hidden Bear in Toblerone
By
–
-

Asymmetric Transformer Architecture with Masked Patch Reconstruction
By
–
Ours is an asymmetric transformer encoder-decoder architecture: encoder on unmasked patches + decoder on full patches. To promote learning, we add an auxiliary task of reconstructing masked patches to denoising score-matching objective that learns the score of unmasked patches.
-
Deploy Computer Vision Apps with Hugging Face Spaces and Gradio
By
–
🎥 New Video Alert!https://t.co/PkU9Felksh
— Satya Mallick (@LearnOpenCV) 19 juin 2023
Unlock the secrets of deploying a Computer Vision App in our latest video!
🚀We simplify the process with Hugging Face Spaces and Gradio. Tune in to elevate your #techgame! #ComputerVision #AppDeployment #HuggingFace #Gradio #ai… pic.twitter.com/aKydWO9VGiNew Alert! https://
youtube.com/watch?v=6b3S2D
2TiAo
… Unlock the secrets of deploying a Computer Vision App in our latest video! We simplify the process with Hugging Face Spaces and Gradio. Tune in to elevate your #techgame! #ComputerVision #AppDeployment #HuggingFace #Gradio #ai -

MagicBrush: Manually Annotated Dataset for Instruction-Guided Image Editing
By
–
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
— AK (@_akhaliq) 19 juin 2023
paper page: https://t.co/T6N8UmgEdz
Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are… pic.twitter.com/gu3PCBokeuMagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing paper page: https://
huggingface.co/papers/2306.10
012
… Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are -

OCTScenes: Real-World Dataset for Object-Centric Learning
By
–
OCTScenes: A Versatile Real-World Dataset of Tabletop Scenes for Object-Centric Learning paper page: https://
huggingface.co/papers/2306.09
682
… Humans possess the cognitive ability to comprehend scenes in a compositional manner. To empower AI systems with similar abilities, object-centric -

Robot Learning with Sensorimotor Pre-training Using Transformer
By
–
Robot Learning with Sensorimotor Pre-training
— AK (@_akhaliq) 19 juin 2023
paper page: https://t.co/JNC4XuEu2f
present a self-supervised sensorimotor pre-training approach for robotics. Our model, called RPT, is a Transformer that operates on sequences of sensorimotor tokens. Given a sequence of camera… pic.twitter.com/DdW8qWTcDPRobot Learning with Sensorimotor Pre-training paper page: https://
huggingface.co/papers/2306.10
007
… present a self-supervised sensorimotor pre-training approach for robotics. Our model, called RPT, is a Transformer that operates on sequences of sensorimotor tokens. Given a sequence of camera -

AvatarBooth: High-Quality Customizable 3D Human Avatar Generation
By
–
AvatarBooth: High-Quality and Customizable 3D Human Avatar Generation
— AK (@_akhaliq) 19 juin 2023
paper page: https://t.co/5WIusjAJDA
introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize… pic.twitter.com/zzfmvxsirbAvatarBooth: High-Quality and Customizable 3D Human Avatar Generation paper page: https://
huggingface.co/papers/2306.09
864
… introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize -

CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos
By
–
CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models paper page: https://
huggingface.co/papers/2306.09
635
… Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text -

Scaling Open-Vocabulary Object Detection with Vision-Language Models
By
–
Scaling Open-Vocabulary Object Detection paper page: https://
huggingface.co/papers/2306.09
683
… Open-vocabulary object detection has benefited greatly from pretrained vision-language models, but is still limited by the amount of available detection training data. While detection training data can -
Computer Vision Gems from Niantic Labs
By
–
@eric_brachmann @dantkz lots of computer vision gems @NianticLabs