Check out MusicGen Stereo in our Transformers Library! 🎶🤗 https://t.co/7sCtsAWg6q
— Hugging Face (@huggingface) 10 novembre 2023
Check out MusicGen Stereo in our Transformers Library!
By
–
Check out MusicGen Stereo in our Transformers Library! 🎶🤗 https://t.co/7sCtsAWg6q
— Hugging Face (@huggingface) 10 novembre 2023
Check out MusicGen Stereo in our Transformers Library!
By
–
Working on a demo of the new @OpenAI Assistants API interface via text message. Feels extremely natural to send requests via text and means you can be multiplayer + multi-modal natively. No need to teach someone a new interface. This is the future.

By
–
RAG with multi-modal embeddings The launch of the GPT-4V API has sparked a surge in interest for multi-modal RAG applications. Multi-modal embeddings will become central to these apps, mapping text and images into a common space for retrieval. This week, @trychroma rolled
By
–
Exploring real time LCM and my camera feed
— fofr (@fofrAI) 10 novembre 2023
👁️🔦 pic.twitter.com/e7HserAVUP
Exploring real time LCM and my camera feed
By
–
GPT-5, codenamed 'Gobi,' rumored for early 2024 release, will handle video interpretation as part of its multimodality. Altman says 'What we launched [at Dev Day] is going to look very quaint relative to what we're busy creating for you now.' '[GPT-5] will work for most things.
By
–
Have you used the speech to text on ChatGPT iOS. It works in the middle of a stadium or with you covering your mouth with a pillow I assume this will work equally great
By
–
Incredibly proud of the @OpenAI team for all the product launches on DevDay. Here’s a roundup of what’s new for developers: GPT-4 Turbo with 128K context and lower prices, the new Assistants API, Vision capabilities, DALL·E 3 API, TTS API, and more! https://
openai.com/blog/new-model
s-and-developer-products-announced-at-devday
…
By
–
This is a cool project with a number of interesting ideas in it. I took a look at the paper to break down how it works:
— Aaron Ng (@localghost) 9 novembre 2023
1. The user is shown a screen with flickering objects.
2. They focus on a flickering object for 10s to confirm.
3. They clench their jaws to reject.
4. They… https://t.co/3yQJbcPRt5
This is a cool project with a number of interesting ideas in it. I took a look at the paper to break down how it works: 1. The user is shown a screen with flickering objects.
2. They focus on a flickering object for 10s to confirm.
3. They clench their jaws to reject.
4. They

By
–
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ye et al.: https://
arxiv.org/abs/2311.04257 #ArtificialIntelligence #DeepLearning #MachineLearning

By
–
Multi-modal LangSmith Tracing With OpenAI releasing GPT4-V, we've now adding support for images in LangSmith tracing! This means you can now see images as they appear in your LLM calls, making it easy to debug multi-modal RAG pipelines Example Trace: