Let me summarize: – users don’t tend to generate 100 codebases or videos and pick the one that is perfect for them. They want fine-grained edit capability. – users tend to generate 100 images or 100 text pieces and pick nearest match to what they want.
MULTIMODAL AI
-
Generating Infinite Variants: Why Code and Video Differ from Images
By
–
With images and text, it seems perfectly reasonable to generate infinite variants till you’re happy. Code and video require too much manipulation of generations if you want changes and that requires skill.
-
Image generation versus code video control trade-offs
By
–
With images, the end users can obviously pull generations into photoshop and edit, but they tend to very often just generate *MORE* variations and just pick 1 Basically with code and video nobody wants to sacrifice control and it doesn’t make sense to generate infinite variation
-
ImageAI Differs from CodeAI and VideoAI in Usage Requirements
By
–
ImageAI is different in the way it is used versus codeAI or videoAI. If you want to modify anything that code AI generates, you need to understand code. Same w video. The output generated will never be perfect and you will want to edit some files (code) or frames (video). More
-

AVFormer Achieves State-of-the-Art Audiovisual Speech Recognition
By
–
Presenting AVFormer, a simple method for injecting visual information into frozen speech models for zero-shot audiovisual (AV) automatic speech recognition (ASR). Read about how AVFormer achieves state-of-the-art AV-ASR performance and more → https://
goo.gle/3IU40P3 -

AI Capability Doubling: Implications for Next 1-2 Years
By
–
These seem like reasonable assumptions with current AI tech. What happens if AI doubles in its capability / reliability over the next 1 – 2 years?
-
Inpainting technique demonstrates advanced generative AI capabilities
By
–
Now that’s in-painting if I’ve seen it
-

ObjectFolder Benchmark: Multisensory Learning Neural Real Objects
By
–
The ObjectFolder Benchmark: Multisensory Learning with Neural and Real Objects paper page: https://
huggingface.co/papers/2306.00
956
… introduce the ObjectFolder Benchmark, a benchmark suite of 10 tasks for multisensory object-centric learning, centered around object recognition, reconstruction, -
GenMM: Generative Motion Matching from Example Sequences
By
–
Example-based Motion Synthesis via Generative Motion Matching
— AK (@_akhaliq) 2 juin 2023
paper page: https://t.co/3w0A1i8wWS
present GenMM, a generative model that "mines" as many diverse motions as possible from a single or few example sequences. In stark contrast to existing data-driven methods, which… pic.twitter.com/bS0rGVCJJ1Example-based Motion Synthesis via Generative Motion Matching paper page: https://
huggingface.co/papers/2306.00
378
… present GenMM, a generative model that "mines" as many diverse motions as possible from a single or few example sequences. In stark contrast to existing data-driven methods, which -

StableRep: Learning Visual Representations from Synthetic Text-to-Image Models
By
–
StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners paper page: https://
huggingface.co/papers/2306.00
984
… We investigate the potential of learning visual representations using synthetic images generated by text-to-image models. This is a natural