It's been a another wild month in AI & Deep Learning research. I curated and summarized noteworthy papers here: https://
magazine.sebastianraschka.com/p/ai-research-
highlights-in-3-sentences-2a1/
… Ranging from new optimizers for LLMs to new scaling laws for vision transformers.
MULTIMODAL AI
-
AI Research Highlights: LLM Optimizers and Vision Transformer Scaling Laws
By
–
-
Otter: Multi-modal Model with Enhanced Instruction Following
By
–
Otter is a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo developed by @laion_ai
), trained on MIMIC-IT dataset and showcasing improved instruction-following and in-context learning ability. -
Killer App for Multi-Modal Foundation Models Finally Arrives
By
–
Finally, a killer app for multi-modal foundation models! 🀄️🤖 https://t.co/t9w9PSjyDl
— hardmaru (@hardmaru) 10 juin 2023Finally, a killer app for multi-modal foundation models!
-
New short: generate music with AI and text (MusicGEN)
By
–
New YouTube short! It is possible to generate music using #AI with just a text and it’s amazing!! Less than 48 hours ago, @MetaAI released a new paper with the code to make #MusicGEN work. I explain it to you.
-

Imagen Editor: State-of-the-Art Text-Guided Image Editing Model
By
–
Imagen Editor is a state-of-the-art model for text-guided image editing of generated and photographed visuals. Learn how we evaluate Imagen Editor using EditBench, a new benchmark that gauges image-text alignment quality across multiple dimensions → https://
goo.gle/3P0KOmv -
Building Media Company with Apple Vision Pro and AI
By
–
The $450 a month that I'm getting from subscriptions will pay for an Apple Vision Pro and I am building a media company around that and AI. I find that having a few people who believe in me and what I do enough to kick in a couple of bucks does give me mental confidence and is
-
Nvidia releases ATT3D for text-to-3D object synthesis
By
–
Nvidia just released ATT3D: Amortized Text-To-3D Object Synthesis
— AK (@_akhaliq) 9 juin 2023
project page: https://t.co/3yAK3Hh4II
Text-to-3D modeling has seen exciting progress by combining generative text-to-image models with image-to-3D methods like Neural Radiance Fields. DreamFusion recently… pic.twitter.com/tOvP31EFS0Nvidia just released ATT3D: Amortized Text-To-3D Object Synthesis project page: https://
research.nvidia.com/labs/toronto-a
i/ATT3D/
… Text-to-3D modeling has seen exciting progress by combining generative text-to-image models with image-to-3D methods like Neural Radiance Fields. DreamFusion recently -
V 1.5 Image Generation Timeline Details
By
–
V 1.5 (they generated the images between late last year and early this year… these and a bunch of other details are in a section at the end of the piece).
-
MusicGen: AI-Generated Music for Meditation
By
–
🧘♀️Meditate with an AI-generated melody ☮️
— Vaibhav (VB) Srivastav (@reach_vb) 9 juin 2023
Brought to you by, MusicGen – A simple and controllable music generation model by @MetaAI🎶
Models on the🤗Hub: https://t.co/BXQFj3nZ7q
Check it out here 👉 https://t.co/uNo3mLVMIu pic.twitter.com/Jsw9yRuhe2Meditate with an AI-generated melody Brought to you by, MusicGen – A simple and controllable music generation model by @MetaAI Models on theHub: https://
huggingface.co/facebook/music
gen-large
… Check it out here https://
huggingface.co/spaces/faceboo
k/MusicGen
… -
Meta Releases MusicGen: Controllable Music Generation Model
By
–
Meta just released MusicGen, a simple and controllable model for music generation
— AK (@_akhaliq) 9 juin 2023
MusicGen is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods like MusicLM, MusicGen doesn't not… pic.twitter.com/kFCOrAmLShMeta just released MusicGen, a simple and controllable model for music generation MusicGen is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods like MusicLM, MusicGen doesn't not