New (2nd edition) from @PacktDataML available at http://
amzn.to/4tULP1b RAG-Driven Generative AI โ Build MAS-RAG with DualRAG, GraphRAG, multimodal video pipelines, and Oracle Database 23ai ๐๐ฒ๐ ๐๐ฒ๐ฎ๐๐๐ฟ๐ฒ๐:
Master DualRAG by combining vector search with SQL filtering
MULTIMODAL AI
-

RAG-Driven Generative AI 2nd Edition: Build MAS-RAG with DualRAG
By
–
-
Offline voice assistant on tiny Axelera AI Mini PC
By
–
A full voice assistant, running with no internet connection at all, on a tiny, self-contained device you can put anywhere. This is the new Axelera AI Mini PC running Llama 3.2 1B as the language model, with separate speech-to-text and text-to-speech models alongside it, all onโฆ pic.twitter.com/MUXkSdlJ7P
— Axelera AI (@AxeleraAI) 26 juin 2026A full voice assistant, running with no internet connection at all, on a tiny, self-contained device you can put anywhere. This is the new Axelera AI Mini PC running Llama 3.2 1B as the language model, with separate speech-to-text and text-to-speech models alongside it, all on
-

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
By
–
ViQ Text-Aligned Visual Quantized Representations at Any Resolution
-

Confidence-Aware Tool Orchestration for Robust Video Understanding
By
–
Confidence-Aware Tool Orchestration for Robust Understanding
-
MinerU parses ugly documents into clean Markdown and JSON for LLM workflows
By
–
/5 MinerU parses ugly documents into clean Markdown and JSON for LLM workflows. It supports PDFs, DOCX, PPTX, XLSX, images, and web pages through a VLM + OCR dual engine. It can handle scanned docs, handwriting, formulas to LaTeX, tables to HTML, multi-column layouts, and
-

Frontier AI models fail medical reasoning stress test, study finds
By
–
We stress tested many frontier AI models for multimodal medical reasoning (including GPT-5, Claude 3.5, Gemini 2.5 Pro). Theyโre not ready. Faulty reasoning, use of inappropriate shortcuts, hallucinations. Published today @NatureMedicine https://
nature.com/articles/s4159
1-026-04501-8
โฆ -

HumanEgo: robot learns skills from human egocentric video
By
–
Your robot could learn a new skill just by watching a few minutes of a human wearing smart glasses! University of Maryland presents HumanEgo: a framework that turns 30 minutes of human egocentric video into a zero-shot robot policy. Instead of needing robot data, it extracts
-
Grok’s reports improve by learning from uploaded videos
By
–
Grok's reports are getting better about what it learned by watching the videos I upload:
-
NVIDIA-accelerated AI aids PYLER in brand safety for advertisers
By
–
Every day, millions of videos compete for advertising dollars. Ensuring brands appear alongside the right content requires AI that can understand context at scale.
— NVIDIA (@nvidia) 25 juin 2026
PYLER is helping advertisers improve brand safety and campaign performance with NVIDIA-accelerated AI that analyzesโฆ pic.twitter.com/9xSDjj9e9gEvery day, millions of videos compete for advertising dollars. Ensuring brands appear alongside the right content requires AI that can understand context at scale. PYLER is helping advertisers improve brand safety and campaign performance with NVIDIA-accelerated AI that analyzes
-
Pim de Witte accidentally built the perfect world model data collection business
By
–
on their @latentspacepod we covered how @pimdewitte accidentally made the PERFECT world model data collection business by collecting the world's largest dataset of trainable (video,action) pairs.
— swyx @aiDotEngineer WF (@swyx) 25 juin 2026
turning the attention economy into the attention industry.
congrats Pim!!
linkโฆ https://t.co/Al4NGdX71W pic.twitter.com/opgI2Kk95Kon their @latentspacepod we covered how @pimdewitte accidentally made the PERFECT world model data collection business by collecting the world's largest dataset of trainable (video,action) pairs. turning the attention economy into the attention industry. congrats Pim!! link