The Luma team is absolutely killing it
MULTIMODAL AI
-
Text-to-3D Generation Creates Full Scene Assets
By
–
I'm genuinely blown away by this.
— Linus ✦ Ekenstam (@LinusEkenstam) 8 novembre 2023
The leap from text descriptions straight to 3D models? It's next-level.
Think about the possibility: a stream of prompts turns into a treasure trove of 3D pieces. Gather them, and you've got a full scene ready to come to life.
The thought… pic.twitter.com/x79WEeY1iqI'm genuinely blown away by this. The leap from text descriptions straight to 3D models? It's next-level. Think about the possibility: a stream of prompts turns into a treasure trove of 3D pieces. Gather them, and you've got a full scene ready to come to life. The thought
-

GPT-4 Vision and DALL-E 3 iterative avatar generation
By
–
I'm using GPT4 vision and dalle3 together to try and reproduce my abstract avatar. I asked it to: – analyze the original image
– analyze the generation
– iterate the prompt to improve accuracy These were the results of 4 iterations. -
Voice Clone Partnership with ElevenLabs Announced
By
–
Yeah the voice clone too, they partnered with Elevenlabs.
-

OtterHD: High-Resolution Multi-modality AI Model
By
–
OtterHD: A High-Resolution Multi-modality Model Li et al.: https://
arxiv.org/abs/2311.04219 #ArtificialIntelligence #DeepLearning #MachineLearning -
OpenAI Releases Text-to-Speech API with Six Natural Voices
By
–
We released the text-to-speech API yesterday at OpenAI DevDay! It offers six preset voices to generate incredibly natural-sounding audio: https://
platform.openai.com/docs/guides/te
xt-to-speech
… -
OpenAI Text-to-Speech API: Weird Experiments Thread
By
–
OpenAI text2speech API.
— fofr (@fofrAI) 7 novembre 2023
A thread of weird experiments 🧵
Based. pic.twitter.com/yseNPDeSiXOpenAI text2speech API. A thread of weird experiments Based.
-
GPT-4 Turbo: Biggest ChatGPT Update Yet
By
–
Biggest ChatGPT Update Yet: GPT-4 Turbo has landed! Get the full story https://
godofprompt.ai/blog/gpt-4-tur
bo-the-biggest-update-for-chatgpt
… Check out these game-changing features: – 128k context window – Updated to April 2023 – 'GPT Builder' for custom chatbots – Multimodal AI for visual learners – -
Real-time translation unlocking 8 billion connected nodes globally
By
–
Correct, about 18% speaks English, at various levels. If we could have true real-time translation between languages, we would essentially make the network power of the world 8 billion nodes strong. Information could flow naturally between all 8bn nodes.
-
Model Performance on Low Quality Video Input
By
–
I think the model held up well here, there are some instances with flickers, but given the low quality input video I'm surprised.