This is also our "v1". Even though the tool is something we use everyday, we are going to keep getting better. Remember Midjourney's March 2022 version? Look at where they are now. Just like our evolution from the last model and approach to this new one, we will keep improving!
MULTIMODAL AI
-

Vision Transformers Interpretation Tutorial at CVPR 2023
By
–
Catch our very own @RisingSayak in collaboration with @hila_chefer at #CVPR2023 as they deliver their tutorial — "All Things ViTs" presenting methods to explain & intepret Vision Transformers! They'll have @MokadyRon as a guest speaker Website https://
all-things-vits.github.io/atv -
ATT3D: Advanced Text-to-3D Object Synthesis Technology
By
–
ATT3D: Amortized Text-to-3D Object Synthesis
— AK (@_akhaliq) 14 juin 2023
paper page: https://t.co/jINaEFQvWR
Text-to-3D modelling has seen exciting progress by combining generative text-to-image models with image-to-3D methods like Neural Radiance Fields. DreamFusion recently achieved high-quality results… pic.twitter.com/uFMHgImJCEATT3D: Amortized Text-to-3D Object Synthesis paper page: https://
huggingface.co/papers/2306.07
349
… Text-to-3D modelling has seen exciting progress by combining generative text-to-image models with image-to-3D methods like Neural Radiance Fields. DreamFusion recently achieved high-quality results -
Neural Scene Chronology: Time-Varying 3D Reconstruction from Internet Photos
By
–
Neural Scene Chronology
— AK (@_akhaliq) 14 juin 2023
paper page: https://t.co/oUdnOhTYls
In this work, we aim to reconstruct a time-varying 3D model, capable of rendering photo-realistic renderings with independent control of viewpoint, illumination, and time, from Internet photos of large-scale landmarks.… pic.twitter.com/pmIuSzH6LNNeural Scene Chronology paper page: https://
huggingface.co/papers/2306.07
970
… In this work, we aim to reconstruct a time-varying 3D model, capable of rendering photo-realistic renderings with independent control of viewpoint, illumination, and time, from Internet photos of large-scale landmarks. -

GeneCIS: Benchmark for General Conditional Image Similarity
By
–
GeneCIS: A Benchmark for General Conditional Image Similarity paper page: https://
huggingface.co/papers/2306.07
969
… argue that there are many notions of 'similarity' and that models, like humans, should be able to adapt to these dynamically. This contrasts with most representation learning -

Image Captioners as Scalable Vision Learning Models
By
–
Image Captioners Are Scalable Vision Learners Too paper page: https://
huggingface.co/papers/2306.07
915
… Contrastive pretraining on image-text pairs from the web is one of the most popular large-scale pretraining strategies for vision backbones, especially in the context of large multimodal -
Instant Multi-View Head Capture through Learnable Registration
By
–
Instant Multi-View Head Capture through Learnable Registration
— AK (@_akhaliq) 14 juin 2023
paper page: https://t.co/cgIfWuUvki
Existing methods for capturing datasets of 3D heads in dense semantic correspondence are slow, and commonly address the problem in two separate steps; multi-view stereo (MVS)… pic.twitter.com/8h3YoWOHIwInstant Multi-View Head Capture through Learnable Registration paper page: https://
huggingface.co/papers/2306.07
437
… Existing methods for capturing datasets of 3D heads in dense semantic correspondence are slow, and commonly address the problem in two separate steps; multi-view stereo (MVS) -
Rerender A Video: Zero-Shot Text-Guided Video Translation
By
–
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
— AK (@_akhaliq) 14 juin 2023
paper page: https://t.co/Ac6ct1naXx
Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring… pic.twitter.com/Il8xSgsW1kRerender A Video: Zero-Shot Text-Guided Video-to-Video Translation paper page: https://
huggingface.co/papers/2306.07
954
… Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring -

Speech-to-Text Adapter for Enhanced LLM Speech Understanding
By
–
Speech-to-Text Adapter and Speech-to-Entity Retriever Augmented LLMs for Speech Understanding paper page: https://
huggingface.co/papers/2306.07
944
… Large Language Models (LLMs) have been applied in the speech domain, often incurring a performance drop due to misaligned between speech and -
SayTap: Language to Quadrupedal Locomotion Control
By
–
SayTap: Language to Quadrupedal Locomotion
— AK (@_akhaliq) 14 juin 2023
paper page: https://t.co/Dk14Ds1D94
Large language models (LLMs) have demonstrated the potential to perform high-level planning. Yet, it remains a challenge for LLMs to comprehend low-level commands, such as joint angle targets or… pic.twitter.com/BteEUxEmalSayTap: Language to Quadrupedal Locomotion paper page: https://
huggingface.co/papers/2306.07
580
… Large language models (LLMs) have demonstrated the potential to perform high-level planning. Yet, it remains a challenge for LLMs to comprehend low-level commands, such as joint angle targets or
