AI war is escalating Who is winning? 05/12/23 Microsoft Copilot Update:
– GPT4 Turbo + DALLE-3 model upgrades
– Deep Search
– Q&A
– Code interpreter capability 06/12/23 Google Gemini Release:
– Gemini Pro-powered Bard
– Pixel Pro: Auto-summarization
– Pixel Pro:
MULTIMODAL AI
-
Microsoft Copilot vs. Google Gemini: A Comparison of AI Capabilities
By
–
-

Gemini Technical Report: TPU Training, Performance vs GPT Models
By
–
Summary of Gemini's 60-page technical report. 1. Written in Jax and trained using TPUs. The architecture, while not explained in details, seems similar to Flamigo's. 2. Gemini Pro's performance is similar to GPT-3.5 and Gemini Ultra is reported to be better than GPT-4. Nano-1
-
Meta Launches Standalone Image Generator for EMU
By
–
Meta just launched a standalone image generator for EMU too!
-
Multi-modal RAG template unlocks Q&A on slide decks
By
–
⭐️ Multi-modal RAG template ⭐️
— LangChain (@LangChain) 6 décembre 2023
Slide decks are a rich sources of information, but their visual elements are inaccessible to most RAG apps.
Multi-modal LLMs, like GPT-4V, can unlock RAG on slides, enabling Q+A assistants over visual content.
We're releasing a new template to… pic.twitter.com/tLLPjwv8ZcMulti-modal RAG template Slide decks are a rich sources of information, but their visual elements are inaccessible to most RAG apps. Multi-modal LLMs, like GPT-4V, can unlock RAG on slides, enabling Q+A assistants over visual content. We're releasing a new template to
-

Playground-v2 Text-to-Image Model Now Available on Replicate
By
–
Playground-v2 is live on Replicate! It's a new text-to-image model created by the research team at @playground_ai
. Run with an API: https://
replicate.com/playgroundai/p
layground-v2-1024px-aesthetic
… -

Pika 1.0 AI Video Generation Tool Demonstration
By
–
AI video with Pika 1.0 made by @Diesol
— Pika (@pika_labs) 6 décembre 2023
pic.twitter.com/kv4Y60UXNMAI video with Pika 1.0 made by @Diesol
-
Gemini Model Card Analysis Delayed by Jury Duty
By
–
Still reading through the model card/report on Gemini (which I did not make the pre-brief cut for this one, alas) so issue will be coming later due to getting called in for jury duty yesterday!
-

Runway Gen-2: AI Video Generation Tool for Creative Storytelling
By
–
Tell a story in any style with Gen-2.https://t.co/ekldoIshdw pic.twitter.com/iq6IuWHlkz
— Runway (@runwayml) 6 décembre 2023Tell a story in any style with Gen-2. http://
runwayml.com -
Google Launches Gemini: State-of-the-Art Multimodal AI Model
By
–
Exciting times, welcome Gemini (and MMLU>90)! State-of-the-art on 30 out of 32 benchmarks across text, coding, audio, images, and video, with a single model 🤯
— Oriol Vinyals (@OriolVinyalsML) 6 décembre 2023
Co-leading Gemini has been my most exciting endeavor, fueled by a very ambitious goal. And that is just the beginning!… pic.twitter.com/AQ5MvJA4upExciting times, welcome Gemini (and MMLU>90)! State-of-the-art on 30 out of 32 benchmarks across text, coding, audio, images, and video, with a single model Co-leading Gemini has been my most exciting endeavor, fueled by a very ambitious goal. And that is just the beginning!
-
Gemini Multimodal Foundation for AI Agents and Reasoning
By
–
I spoke with @demishassabis ahead of Gemini's launch today. He says the new multimodal model will be a foundation for rapid innovation in software agents, planning and reasoning (a la Q*), gameplay, and even physical robots.