Top ML Papers of the Week (Mar 13 – Mar 19): – GPT-4
– FlexGen
– NeRFMeshing – Resurrecting RNNs
– An Overview of Language Models
– Universal Prompt Retrieval for LLMs
…
MULTIMODAL AI
-
Top ML Papers of the Week: GPT-4 and Language Models
By
–
-
GPT-4 Large Multimodal Model Advances General Knowledge
By
–
1/ GPT-4 – a large multimodal model with broader general knowledge and problem-solving abilities. https://t.co/drO0DXfZK9
— DAIR.AI (@dair_ai) 19 mars 20231/ GPT-4 – a large multimodal model with broader general knowledge and problem-solving abilities.
-
LERF: Language Embeddings in 3D Radiance Fields
By
–
2/ LERF (Language Embedded Radiance Fields) – a method for grounding language embeddings from models like CLIP into NeRF; this enables open-ended language queries in 3D.https://t.co/0zbkWzzTUT
— DAIR.AI (@dair_ai) 19 mars 20232/ LERF (Language Embedded Radiance Fields) – a method for grounding language embeddings from models like CLIP into NeRF; this enables open-ended language queries in 3D.
-
Future of Movie Storyboarding with Photorealistic AI Concepts
By
–
Storyboarding in movies will be so wild in 5 years. You basically have photorealistic concepts of each shot. This combined with digital walls will be interesting
-
Optimizing Pepper Voice Parameters in Noisy Environments
By
–
@EmileAndHisBots A quick and dirty takeaway for you is to consider increasing Pepper's pitch in loud environments. Joyful style works, although if saying something negative you can consider using a Lombard voice with didactic (to slow down speed) and increasing pitch by +130.
-

Cohere Multilingual Models Power Interactive Movie Discovery App
By
–
3/ With the app, you'll use Cohere's multilingual models to embed movie descriptions into language-invariant embeddings. Say hello to an interactive & user-friendly interface that makes finding the perfect movie in your preferred language a breeze!
-

ControlNet Models Surge on Hugging Face Community
By
–
So cool to see so much controlnet stuff on HF: https://
huggingface.co/search/full-te
xt?q=controlnet
…! Congrats @lvminzhang @magrawala @Stanford -

How Self-Driving Cars Perceive and Navigate the World
By
–
This is what the world looks like to a self-driving car. @CNET v/ @CurieuxExplorer #AutonomousVehicles #SelfDrivingCars #electriccars #ElectricVehicles #Automotive #5G @AlbertoEMachado @KanezaDiane @asokan_telecom @debashis_dutta @guidaautonoma @bimedotcom @Shi4Tech
-
Vid2Seq: Visual Language Model for Dense Video Captioning
By
–
Introducing Vid2Seq, a visual language model for dense video captioning that simply predicts all event boundaries and captions as a single sequence of tokens. Learn more about how it achieves state-of-the-art results on various benchmarks → https://t.co/CgQXVBnNYs pic.twitter.com/oQ1fXwEBAj
— Google AI (@GoogleAI) 17 mars 2023Introducing Vid2Seq, a visual language model for dense video captioning that simply predicts all event boundaries and captions as a single sequence of tokens. Learn more about how it achieves state-of-the-art results on various benchmarks → https://
goo.gle/3JNz97v -

Real-Time Sign Language Translation Using Artificial Intelligence
By
–
Translating Sign Language in Real Time
— Ronald van Loon (@Ronald_vanLoon) 17 mars 2023
via @pascal_bornet#AI #ArtificialIntelligence #Innovation #MachineLeaning #TechForGood
cc: @dirkschaar @space_mog @maxjcm @ronald_vanloon pic.twitter.com/d19jANtgySTranslating Sign Language in Real Time via @pascal_bornet #AI #ArtificialIntelligence #Innovation #MachineLeaning #TechForGood cc: @dirkschaar @space_mog @maxjcm @ronald_vanloon