AudioCraft by Meta AI is a one-stop codebase for generative audio consisting of three models: MusicGen, AudioGen & EnCodec — supporting both compression + generation of high-quality music and sound effects from text.. Get the code
MULTIMODAL AI
-

Gen-2 Video Generation: Creating 18-Second Content
By
–
Learn how to generate up to 18 seconds with Gen-2. pic.twitter.com/CLMrpFT9w9
— Runway (@runwayml) 11 août 2023Learn how to generate up to 18 seconds with Gen-2.
-
Continuous prediction learning from webcam visual environment data
By
–
If you walk around with a webcam you can be constantly predicting what you'll see next, and checking if those predictions match reality. So the environment is generating data to learn from, but not humans as such. Open environment, open-ended data gathering.
-

Multimodal Medical AI: Advancing Healthcare with Machine Learning
By
–
Multimodal medical AI https://
bit.ly/45ghLAf #AI #MachineLearning #DeepLearning #LLMs #DataScience -

Limited Time Offer: Computer Vision AI Courses Available Now
By
–
Seize the moment before the price goes up! https://
opencv.org/university/ #AI #course #computervision -

Digital Surgery: AI and Robotics Transform Healthcare Globally
By
–
Digital surgery, at the intersection of AI, robotics, augmented reality, the IoT, and real-time data analytics, is seen as the next evolution in surgical procedures. It also has the potential to reduce health disparities in developing countries. Microblog @antgrasso
-
Pro Dashboard for AI VTuber Creators and Avatar Management
By
–
Introducing a pro dashboard for Hyper AI Creators and VTubers to manage:
— Aaron Ng (@localghost) 9 août 2023
+ VRM Avatars
+ Backgrounds
+ Props
and everything else they need to build great stories on the app. pic.twitter.com/1w9SOhOsV6Introducing a pro dashboard for Hyper AI Creators and VTubers to manage: + VRM Avatars
+ Backgrounds
+ Props and everything else they need to build great stories on the app. -
VRDU: Visual Rich Document Understanding Benchmark Dataset
By
–
Presenting Visually Rich Document Understanding (VRDU), a dataset for better tracking of document understanding task progress. Read about VRDU and the five requirements needed to create benchmarks that capture the complexity of real-world applications → https://t.co/YQsKii17WD pic.twitter.com/oaGB3yzxwX
— Google AI (@GoogleAI) 9 août 2023Presenting Visually Rich Document Understanding (VRDU), a dataset for better tracking of document understanding task progress. Read about VRDU and the five requirements needed to create benchmarks that capture the complexity of real-world applications → https://
bit.ly/3KApkJU -

SDXL Generates GTA V Style Images via Replicate
By
–
3. GTA V by @PontusAurdal https://
replicate.com/pwntus/sdxl-gt
a-v
…