Rly cool. One of the rarest videos that goes beyond animating realistic images and focuses on the style reference
GENERATIVE AI
-
Multimodal AI and audio generation capabilities
By
–
I only heard that 4o multimodal is capable of doing it. Now maybe a model behind NotebookLM audio overviews as well. Haven’t checked 11labs enough to say if the model supports it or not tbh. But it will be a huge feature for audiobooks for sure
-

Refik Anadol discusses Llama in Large Nature Model project
By
–
On the latest episode of the Boz To The Future podcast, media artist and director, @RefikAnadol shared how his studio used Llama as part of "Large Nature Model: A Living Archive". Read more on the project and watch the whole conversation https://
go.fb.me/pfhg9i -
Salesforce Venture invests in AI startups
By
–
Some more wonga for AI startups (aka prospective recipients of big, big bags of ) from corporate venture coffers – this time, from Salesforce's venture arm:
-

AI Engineer: Running and Hosting LLM-Generated Code for Business
By
–
LLMs can write code to automate business processes or create chatbots/AI agents, but how would you run this code? Our AI engineer can host code, run pipelines, connect to hundreds of apps, and understand structured and unstructured data.
-
Comparison of AI Voice Features: Gemini Live vs. ChatGPT
By
–
Gemini Live is an equivalent of ChatGPT standard voice that has been there for a long time It didn’t get Advanced mode features yet (there are planned soon though)
-
Comparative Analysis of AI Voice Assistant Capabilities
By
–
The quality of the product is quite good. Better voice experience than Gemini Live for example. ChatGPT voice is more useful cuz of its tools – web access for example. It cannot generate sounds but supports interruptions and emotion detection. You should treat is as a demo
-

AI Compute Power: 10,000x GPT-4 by 2030, Is It a Bottleneck?
By
–
Analysis: By 2030, we will see a 10,000-fold increase in computing power compared to GPT-4. Epoch AI's analysis forecasts that AI training runs could reach up to 2e29 FLOP (floating point operations) in about 5 years. Will we hit an electricity/chips/data bottleneck? Link to
-

Diffusion Approach to Radiance Field Relighting with Multi-Illumination
By
–
A Diffusion Approach to Radiance Field Relighting using Multi-Illumination Synthesis
— AK (@_akhaliq) 16 septembre 2024
discuss: https://t.co/2k0Bw3j5nG
Relighting radiance fields is severely underconstrained for multi-view data, which is most often captured under a single illumination condition; It is… pic.twitter.com/BD1HcvEpX0A Diffusion Approach to Radiance Field Relighting using Multi-Illumination Synthesis discuss: https://
huggingface.co/papers/2409.08
947
… Relighting radiance fields is severely underconstrained for multi-view data, which is most often captured under a single illumination condition; It is -

InstantDrag: Improving Interactivity in Drag-Based Image Editing
By
–
InstantDrag
— AK (@_akhaliq) 16 septembre 2024
Improving Interactivity in Drag-based Image Editing
discuss: https://t.co/e0oxkJpWao
Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a… pic.twitter.com/QwAvTbPB5LInstantDrag Improving Interactivity in Drag-based Image Editing discuss: https://
huggingface.co/papers/2409.08
857
… Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a