bet, i was thinking there was some computer vision edge detection going on in the images themselves — do we know what gen 3 video to video uses? depth + rgb + other stuff?
GENERATIVE AI
-
How AI Access Points Shape Model Behavior and Performance
By
–
Does how you access AI matter? Absolutely. The Gemini model inside of the Gemini chat app is different from the Gemini model inside AI Studio—and for good reason, they’re trying to do two different things. Upstream decisions can make a huge impact.
-

Pangea: Open Multilingual Multimodal LLM for 39 Languages
By
–
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
—————————–
Paper: https://
arxiv.org/abs/2410.16153
Project website: https://
neulab.github.io/Pangea/
Model: https://
huggingface.co/neulab/Pangea-
7B
…
Extended post by @xiangyue96 : -

CMU releases NaturalBench for multimodal LLM evaluation
By
–
This week, we(two amazing research groups at CMU) released two projects in multimodal evaluation and multilingual multimodal LLMs. – NaturalBench: A vision-centric VQA evaluation benchmark that challenges all leading multimodal LLMs while using natural images and simple
-
Open Source AI Project Code Available for Experimentation
By
–
But, it’s open – with all the code open source, you can fiddle with it, try different prompts, things etc The final output is not at that level, but, it only gets better from here,
-
F5-TTS Audio Quality Optimization with Reference Speakers
By
–
Nice! Lmk how it goes, I’ll be back in office tomorrow will take a deeper look. Btw from my experience w/ F5-TTS – the generation quality depends quite a bit on the reference audio – might be worth checking with different speaker prompts.
-
OpenAI o1 Model Reasoning Patterns and Performance Analysis
By
–
10). Reasoning Patterns of OpenAI’s o1 Model – when compared with other test-time compute methods, o1 achieved the best performance across most datasets; the authors observe that the most commonly used reasoning patterns in o1 are divide and conquer and self-refinement.
-
SynthID-Text: Scalable Watermarking Scheme for LLMs
By
–
9). Scalable Watermarking for LLMs – proposes SynthID-Text, a text-watermarking scheme that can preserve text quality in LLMs, enable high detection accuracy, and minimize latency overhead…https://t.co/isvnHcP610
— DAIR.AI (@dair_ai) 27 octobre 20249). Scalable Watermarking for LLMs – proposes SynthID-Text, a text-watermarking scheme that can preserve text quality in LLMs, enable high detection accuracy, and minimize latency overhead…
-
Feature Steering in LLMs: Evaluating Social Bias Control
By
–
6). Evaluation Feature Steering in LLMs – evaluates featuring steering in LLMs using an experiment that artificially dials up and down various features to analyze changes in model outputs; it focused on 29 features related to social biases and study if feature steering can help
-

IBM Granite 3.0 Lightweight Foundation Models for Enterprise
By
–
7). Granite 3.0 – presents lightweight foundation models ranging from 400 million to 8B parameters; supports coding, RAG, reasoning, and function calling, focusing on enterprise use cases, including on-premise and on-device settings.