Google released a set of TranslateGemma models in 4B, 12B, and 27B parameter sizes, with support for 55 languages and a lower error rate.
MULTIMODAL AI
-
Interactive 3D Models for Robot Learning in Dynamic Environments
By
–
Interactive 3D world model is a highly intuitive representation for learning robotics actions in dynamic and complex environments. Here is our most recent work on this 🤖 https://t.co/si6T0gyhrf
— Fei-Fei Li (@drfeifei) 15 janvier 2026Interactive 3D world model is a highly intuitive representation for learning robotics actions in dynamic and complex environments. Here is our most recent work on this
-
AI Character Entertainment Assistant for Personal Use
By
–
i’ve been entertaining my guests all day, now I need my Character to entertain me. pic.twitter.com/39Qwm5z26T
— Character.AI (@character_ai) 15 janvier 2026i’ve been entertaining my guests all day, now I need my Character to entertain me.
-
FLUX.2 Klein: Fast Compact Image Generation Model Released
By
–
FLUX.2 [klein] is here A compact, lightning-fast image model from @bfl_ml Generate 1MP images in ~500ms or 4MP in under 2s. Supports image-to-image and precise editing with up to 5 input images.
-

Gemini’s Personal Intelligence and Latest AI Tools Update
By
–
Top stories in AI today: – Gemini’s ‘Personal Intelligence’ shift
– McConaughey trademarks himself against deepfakes
– Get most out of Gemini in Gmail
– Z AI’s image model trained on Huawei chips
– 4 new AI tools, community workflows, and more Read more: https://
therundown.ai/p/geminis-pers
onal-intelligence-upgrade
… -
Computer Vision and Agentic AI Transform Video Analytics
By
–
Computer vision can see what happened — agentic AI explains why it matters and what to do next.
— NVIDIA AI (@NVIDIAAI) 14 janvier 2026
Here’s how teams are upgrading video analytics with vision language models:
🔍 Turn video into searchable intelligence
🧠 Add context and reasoning to system alerts
⚡ Summarize… pic.twitter.com/vWuUwQX6IaComputer vision can see what happened — agentic AI explains why it matters and what to do next. Here’s how teams are upgrading video analytics with vision language models: Turn video into searchable intelligence Add context and reasoning to system alerts Summarize
-

Salesforce joins Hugging Face as major enterprise customer
By
–
Excited to welcome @salesforce as our latest enterprise customer! Already massive contributions (180 public models like Blip) and can't wait for what they'll do next! https://
huggingface.co/Salesforce -

Google Adds Personal Data to Gemini
By
–

Google introduced Personal Intelligence which can allow Gemini to use your data from Google Photos, Gmail, YouTube and Google Search as a context. Seems like we will see Google Photos attachment option in Gemini pretty soon!
-

New Document AI Course: OCR to Agentic Document Extraction
By
–
New course: Document AI: From OCR to Agentic Doc Extraction, built with @LandingAI, where I'm executive chairman, and taught by David Park and Andrea Kropp.
— Andrew Ng (@AndrewYNg) 14 janvier 2026
Much of the world's data is locked in PDFs, JPEGs, and other documents. This short course shows you how to build agentic… pic.twitter.com/dG9SwmFgKqNew course: Document AI: From OCR to Agentic Doc Extraction, built with @LandingAI, where I'm executive chairman, and taught by David Park and Andrea Kropp. Much of the world's data is locked in PDFs, JPEGs, and other documents. This short course shows you how to build agentic workflows that process documents accurately: breaking them into parts, examining each piece carefully, and extracting information through multiple iterations. Traditional Optical Character Recognition (OCR) captures text but loses context from table headers, chart captions, or reading order of columns. After exploring OCR's limitations, you’ll use LandingAI's Agentic Document Extraction (ADE) framework to process documents. ADE treats pages as visually — as images — to parse information and extract fields. Skills you'll gain: – Build agents to convert unstructured files into structured Markdown/HTML and JSON – Use ADE to parse complex data like forms, handwriting, or equations – Map extracted information to named fields using a specified schema, with bounding boxes for grounding and validation – Deploy RAG applications with event-driven document processing Come learn about the best tools for processing documents like financial invoices, medical records, or academic papers intelligently: deeplearning.ai/short-course…
→ View original post on X — @andrewyng, 2026-01-14 17:42 UTC
-

Agentic Hackathon Showcases Impressive AI Agent Projects
By
–
Super impressed by the projects at the Agentic Hackathon last weekend! Many teams work on really hard/important problems: * Long running tasks: memory management, recovering from mid-task failures, and maintaining consistency across steps and sub-agents * Adaptive retrieval from multiple sources: databases, search indices, and websites * Agents that work with voice, video, and even 3D environments If you are in SF, come check out the finalist demos tomorrow! luma.com/6bd4bt9j There will be talks by Douglas Eck, who is doing amazing work with Veo and Imagen and many other awesome folks. Thanks @MongoDB and @cerebral_valley for hosting and for letting me serve as a judge for these fantastic projects.
