Microsoft bans US police departments from using enterprise AI tool for facial recognition
@aibreakfast
-

AI language models may predict multiple tokens at once
By
–
The future of AI language models may lie in predicting beyond the next word: Multi-Token Prediction Studies suggest that the human brain predicts multiple words at once when understanding language, utilizing both semantic and syntactic information for broader predictions – now
-
Efficient fine-tuning enables GPT-2 to approach GPT-4 performance
By
–
Most likely explanation for gpt2-chatbot: OpenAI has been working on a more efficient method for fine-tuning language models, and they managed to get GPT-2, a 1.5B parameter model, to perform pretty damn close to GPT-4, which is an order of magnitude larger and more costly to
-
China unveils Vidu, AI video generator rivaling OpenAI’s Sora
By
–
Overnight, China released their own version of OpenAI’s Sora:
— AI Breakfast (@AiBreakfast) 27 avril 2024
AI video generator “Vidu” can create 16 second clips at 1080p:pic.twitter.com/LkACnl1M2oOvernight, China released their own version of OpenAI’s Sora: AI video generator “Vidu” can create 16 second clips at 1080p:
-

Ray-Ban Meta Smart Glasses gain multimodal AI for visual tasks
By
–

It's happening: Ray-Ban Meta Smart Glasses have multimodal AI now: "The primary command is to say “Hey Meta, look and…” You can fill out the rest with phrases like “Tell me what this plant is.” Or read a sign in a different language." Story below ↓
-
GPT-5 Predictions: Multi-modal generation from any input with precise refinement
By
–
Alright, GPT-5 Predictions: Generate text, images, video, 3D assets or music FROM text, images, video, or music (basically it can take anything as input, and produce any medium for output depending on the prompt direction) with much more precise refinement process to massage an
-
Microsoft’s VASA-1: Real-Time Lifelike Audio-Driven Talking Faces
By
–
Microsoft publishes paper on VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
— AI Breakfast (@AiBreakfast) 19 avril 2024
VASA is capable of generating a large spectrum of expressive facial nuances and natural head motions
It can handle long-form audio and stably output seamless talking face videos: pic.twitter.com/FiBb11G1ruMicrosoft publishes paper on VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time VASA is capable of generating a large spectrum of expressive facial nuances and natural head motions It can handle long-form audio and stably output seamless talking face videos:
-
20% of Americans flirt with chatbots: Curiosity and loneliness drive interaction
By
–
Already 20% of Americans have flirted with chatbots, according to this study [linked] Nearly half of them — 47.2% — did so out of curiosity while 23.9% said they were lonely and seeking interactions. It's only April 2024. The line on this graph only goes in one direction from
-

Grok-1.5V becomes multimodal with sample imagery analysis
By
–



Grok goes multimodal with Grok-1.5V Sample imagery analysis:
-
Scalable streaming Transformers for infinite context with bounded resources
By
–
Scalability to infinitely long context: Processes extremely long inputs in a streaming fashion, overcoming limitations of standard Transformers. Bounded memory and compute resources: Achieves high compression ratios while maintaining performance, super cost-effective. Who
