ChatGPT is now in your car. OpenAI has just launched ChatGPT in Apple CarPlay. You speak. ChatGPT responds. Right from your dashboard screen. Hands-free. What's fascinating is the story behind it. A developer, Gui Ferreira, had tweeted in September 2025: "ChatGPT's voice mode
MULTIMODAL AI
-
OpenAI’s GPT-image-2 Revolutionary Advancement in Image Generation
By
–
Holy, OpenAI's GPT-image-2 will crush everything. I remember when everyone laughed at the GPT image because it couldn't generate a proper world map. Those days are over. And even the YouTube image is now indistinguishable from reality. Holy moly.
-
OpenAI’s GPT-Image-2 model leak surpasses Nano Banana Pro
By
–
OpenAI's new image model GPT-Image-2 has leaked It seems to have extremely good world knowledge and great text rendering Possibly better than Nano Banana Pro It's on @arena under code names:
– maskingtape-alpha
– gaffertape-alpha
– packingtape-alpha -
Wan 2.7 Video Generation Model Now Available on Poe
By
–
Wan 2.7 is now live on Poe!
— Poe (@poe_platform) 4 avril 2026
Four video generation modes in one model: Text-to-Video, Image-to-Video, Video Edit, and Reference-to-Video. Supports multi-shot generation, audio-driven output, first-and-last-frame control, and seamless video continuation — from concept to final… pic.twitter.com/f93tAAFB86Wan 2.7 is now live on Poe! Four video generation modes in one model: Text-to-Video, Image-to-Video, Edit, and Reference-to-Video. Supports multi-shot generation, audio-driven output, first-and-last-frame control, and seamless video continuation — from concept to final
-

Intelligent Remote Sensing Agents Transform Earth Observation with AI
By
–
What if Earth observation could truly think for itself? A collaborative team from Hong Kong University of Science and Technology, Northwestern Polytechnical University, Tsinghua University, and international partners have released a seminal survey on "Intelligent Remote Sensing Agents." This new paradigm shows how AI agents integrate perception, planning, memory, and tool execution to autonomously achieve complex geospatial understanding. This transforms remote sensing from passive data collection into proactive, intelligent decision support, far surpassing previous capabilities in urban governance, precision agriculture, ecological monitoring, and emergency response. Intelligent Remote Sensing Agents: A Survey Paper: github.com/PolyX-Research/Aw… Repo: github.com/PolyX-Research/Aw… Our report: mp.weixin.qq.com/s/QYAyTjaAa… 📬 #PapersAccepted by Jiqizhixin
→ View original post on X — @jiqizhixin, 2026-04-04 01:40 UTC
-

Google Gemma 4 31B Open Model Now Live on Poe
By
–
Gemma 4 is now live on Poe. Google's most capable open model: 31B dense parameters, 256K context window, 140+ languages, multimodal inputs, built-in chain-of-thought reasoning, and function calling support. Apache 2.0 licensed. You can try it in Poe app on all platforms and in
-
Gesture Control Technology Powers Advanced Tentacle Robots
By
–
Mind Over Machine: Gesture Control Brings Tentacle Robots to Life
— Ronald van Loon (@Ronald_vanLoon) 4 avril 2026
by @Fabriziobustama
#Robotics #Engineering #ArtificialIntelligence #Innovation #Technology pic.twitter.com/anPcF7WX2hMind Over Machine: Gesture Control Brings Tentacle Robots to Life
by @Fabriziobustama #Robotics #Engineering #ArtificialIntelligence #Innovation #Technology -
Reality Becomes Indistinguishable: The Deepfake Dilemma Unfolds
By
–
It will all seem real. But it is all fake. 🙂
-

Closing the 100,000-Year Robot Data Gap with Code-as-Policy
By
–



Thank your for this excellent summary Junfan! Junfan Zhu 朱俊帆 (@junfanzhu98) How to Close the 100,000-Year Robot “Data Gap” — @Ken_Goldberg (@UCBerkeley) Goldberg’s core claim: end-to-end Vision-Language-Action (VLA) models aren’t delivering. They’re opaque, hard to debug, and fragile under distribution shift. On LIBERO / LIBERO-PRO, models reach ~100% in-distribution, but tiny pose perturbations collapse success to ~17% or 0% (even π₀). This is systematic overfitting, not generalization. Code-as-Policy (CaP) reframes control: LLMs generate executable programs that call structured primitives (perception, 6D pose, motion planning, grasping). Generalization shifts from weights → code. Benefits: interpretable, verifiable, training-free at inference, debuggable. Open question: reliability. CaP-X (arXiv 2603.22435) introduces a full evaluation stack: 🔷 CaP-Gym: unified REPL over RoboSuite + LIBERO-PRO + BEHAVIOR (tabletop → mobile/bimanual, sim→real) 🔷 CaP-Bench: multi-level abstraction tests 🔷 CaP-Agent0: training-free agent (visual differencing, skill library, parallel queries) 🔷 CaP-RL: verifiable reward RL in Python sandbox Results: under perturbations that break VLAs, CaP-X hits ~96% success with strong pose invariance (extreme corners, lighting, object swaps). LIBERO-PRO (50 trials/task): many 100%, lowest ~76–78%. Failures are mostly semantic (label ambiguity), not control. Grasping/planning ≈ solved. CaP-X 2.0 pushes agentic coding: prompt restructuring, failure-analysis primitives, human-in-loop, reusable cloud skill cache. Test-time loop (no retraining): generate → compile → execute → perturb → diagnose → patch. Extensions include Rust backends (reliability) and Graph-as-Policy (GaP) for node-level verification. Core thesis (GOFE + CaP hybrid): pure VLA scaling cannot close the 100,000-year data gap (robot ≈10K hrs vs LLM ≈1.2B hrs). Robotics needs a Good Old-Fashioned Engineering (GOFE) skeleton: modular pipelines, PID (kp, kv), feedforward (e.g., virtual gravity) — inherently pose-invariant. → Build GOFE backbone + CaP brain. Deploy now, collect real data, spin a flywheel to improve modules and future VLAs. Hot 🔥 takes: 🔷 “Robot generalists should get off their high horse.” VLA-only is dogmatic. 🔷 VLAs may win eventually, but near-term progress requires hybrids. 🔷 Reliability (→99.9%) is the real bottleneck, not demos. Comparison 🔷 GOFE: no generality, but available, interpretable, reliable 🔷 VLA: promised generality, but opaque, brittle, not ready 🔷 CaP: generality + available + interpretable; reliability improves via hybrid + iteration Timeline: near-term: structured tasks (declutter, laundry, delivery). ~5 years: major home logistics gains. Full humanoid generalists: far off. Strategy: specialists first + reliability to 99.9%. Bottom line: don’t wait for end-to-end intelligence. Turn LLMs into super-programmers over a GOFE substrate, deploy hybrids, iterate with real data, and asymptotically approach VLA—without burning out the field. 👉🏻More pics: linkedin.com/posts/junfan-zh… — https://nitter.net/junfanzhu98/status/2039953079706247169#m
→ View original post on X — @ken_goldberg, 2026-04-03 23:15 UTC

