GPT-4 can now analyze PDFs A moment of silence for the 10,000 AI startups that just became obsolete.
MULTIMODAL AI
-

CommonCanvas: CC Dataset for Competitive Diffusion Model Training
By
–
8/ CommonCanvas – a dataset of Creative-Commons-licensed (CC) images to train diffusion models competitive with Stable Diffusion 2 (SD2); rely on a transfer learning technique to produce high-quality synthetic captions paired with curated CC images.
-

Matryoshka Diffusion Models for High-Resolution Image Video Synthesis
By
–
3/ Matryoshka Diffusion Models – introduces an end-to-end framework for high-resolution image and video synthesis; involves a diffusion process that denoises inputs at multiple resolutions jointly and uses a NestedUNet architecture…
-

Spectron: End-to-End Spoken Language Model Surpasses Existing Systems
By
–
4/ Spectron – a spoken language model trained end-to-end to directly process spectrograms; it can be fine-tuned to generate high-quality accurate spoken language; surpasses existing spoken language models in speaker preservation and semantic coherence.
-

Spooky Black Light SDXL Fine-Tune Model Released
By
–
Spooky black light SDXL fine tune anyone? https://
replicate.com/fofr/sdxl-blac
k-light
… -
Joint Embedding Predictive Architecture Evolution Since 1993
By
–
World models from video: since 2014.
Learned World Models for planning: since 2018.
Joint Embedding Architectures: since 1993, but a lot more since 2019.
Joint Embedding Predictive Architecture (JEPA) for images: since 2021.
JEPA trained from video: since earlier this year. -
3D-LLM: Bridging Language and 3D World Perception
By
–
With these evolutions come new tools and models such as the 3D-LLM. This novel model is bridging the gap between language and the 3D world around us. It shows potential not just to perceive the world but also act within it.
-
VALL-E Voice Cloning and Drag Your GAN Image Editing Advances
By
–
The new audio system VALL-E can duplicate a voice using a 3-second recording, expanding AI voice imitation limits. Meanwhile, Drag Your Gan applies StyleGAN2 tech for interactive and precise image editing.
-
Deepfakes and AI: Distinguishing Truth from Deception
By
–
Rien dans cette vidéo n'est vrai ! Ceci est un clone de moi même (ou deep-fake) réalisé en quelques secondes à partir d'une IA. Le grand défi des années à venir sera de distinguer le vrai du faux dans un monde saturé de fake news et d'images trompeuses. Soyez vigilants. pic.twitter.com/DPSVgDh0BZ
— Rafik Smati (@RafikSmati) 28 octobre 2023Rien dans cette vidéo n'est vrai ! Ceci est un clone de moi même (ou deep-fake) réalisé en quelques secondes à partir d'une IA. Le grand défi des années à venir sera de distinguer le vrai du faux dans un monde saturé de fake news et d'images trompeuses. Soyez vigilants.

