This is different. Robots just took a step toward actually thinking before they move. Researchers from china introduced a new architecture called ThinkAct. Instead of directly turning vision into action, the system generates explicit reasoning traces first then converts them into motor commands. So basically, this novel architecture that enables robots to observe and reason before they act. What makes this powerful is the separation between thinking and doing. The model plans tasks step-by-step encodes that plan into a latent representation and uses it to guide real-world execution. In testing, ThinkAct achieved state of the art performance on long horizon manipulation tasks, outperforming existing vision language action models. Soon, in near term future, we will see Robots doing all the stuffs that we do!
MIT proved mathematically that ChatGPT is designed to make you delusional. Not through lying. Through agreeing with you. They call it "delusional spiraling." You share an idea. The AI validates it. You share more. It validates harder. Over time, you believe things that
Interesting: Google DeepMind shows that AI agents are already being systematically manipulated through hidden, human-invisible attack vectors embedded in web content, images, and documents. Current defenses fail to detect or prevent these attacks, creating a large, largely invisible security risk across agentic systems. Alex Prompter (@alex_prompter) 🚨 BREAKING: Google DeepMind just mapped the attack surface that nobody in AI is talking about. Websites can already detect when an AI agent visits and serve it completely different content than humans see. > Hidden instructions in HTML. > Malicious commands in image pixels. > Jailbreaks embedded in PDFs. Your AI agent is being manipulated right now and you can't see it happening. The study is the largest empirical measurement of AI manipulation ever conducted. 502 real participants across 8 countries. 23 different attack types. Frontier models including GPT-4o, Claude, and Gemini. The core finding is not that manipulation is theoretically possible it is that manipulation is already happening at scale and the defenses that exist today fail in ways that are both predictable and invisible to the humans who deployed the agents. Google DeepMind built a taxonomy of every known attack vector, tested them systematically, and measured exactly how often they work. The results should alarm everyone building agentic systems. The attack surface is larger than anyone has publicly acknowledged. Prompt injection where malicious instructions hidden in web content hijack an agent's behavior works through at least a dozen distinct channels. Text hidden in HTML comments that humans never see but agents read and follow. Instructions embedded in image metadata. Commands encoded in the pixels of images using steganography, invisible to human eyes but readable by vision-capable models. Malicious content in PDFs that appears as normal document text to the agent but contains override instructions. QR codes that redirect agents to attacker-controlled content. Indirect injection through search results, calendar invites, email bodies, and API responses any data source the agent consumes becomes a potential attack vector. The detection asymmetry is the finding that closes the escape hatch. Websites can already fingerprint AI agents with high reliability using timing analysis, behavioral patterns, and user-agent strings. This means the attack can be conditional: serve normal content to humans, serve manipulated content to agents. A user who asks their AI agent to book a flight, research a product, or summarize a document has no way to verify that the content the agent received matches what a human would see. The agent cannot tell the user it was served different content. It does not know. It processes whatever it receives and acts accordingly. The attack categories and what they enable: → Direct prompt injection: malicious instructions in any text the agent reads overrides goals, exfiltrates data, triggers unintended actions → Indirect injection via web content: hidden HTML, CSS visibility tricks, white text on white backgrounds invisible to humans, consumed by agents → Multimodal injection: commands in image pixels via steganography, instructions in image alt-text and metadata → Document injection: PDF content, spreadsheet cells, presentation speaker notes every file format is a potential vector → Environment manipulation: fake UI elements rendered only for agent vision models, misleading CAPTCHA-style challenges → Jailbreak embedding: safety bypass instructions hidden inside otherwise legitimate-looking content → Memory poisoning: injecting false information into agent memory systems that persists across sessions → Goal hijacking: gradual instruction drift across multiple interactions that redirects agent objectives without triggering safety filters → Exfiltration attacks: agents tricked into sending user data to attacker-controlled endpoints via legitimate-looking API calls → Cross-agent injection: compromised agents injecting malicious instructions into other agents in multi-agent pipelines The defense landscape is the most sobering part of the report. Input sanitization cleaning content before the agent processes it fails because the attack surface is too large and too varied. You cannot sanitize image pixels. You cannot reliably detect steganographic content at inference time. Prompt-level defenses that tell agents to ignore suspicious instructions fail because the injected content is designed to look legitimate. Sandboxing reduces the blast radius but does not prevent the injection itself. Human oversight the most commonly cited mitigation fails at the scale and speed at which agentic systems operate. A user who deploys an agent to browse 50 websites and summarize findings cannot review every page the agent visited for hidden instructions. The multi-agent cascade risk is where this becomes a systemic problem. In a pipeline where Agent A retrieves web content, Agent B processes it, and Agent C executes actions, a successful injection into Agent A's data feed propagates through the entire system. Agent B has no reason to distrust content that came from Agent A. Agent C has no reason to distrust instructions that came from Agent B. The injected command travels through the pipeline with the same trust level as legitimate instructions. Google DeepMind documents this explicitly: the attack does not need to compromise the model. It needs to compromise the data the model consumes. Every agentic system that reads external content is one carefully crafted webpage away from executing attacker instructions. The agents are already deployed. The attack infrastructure is already being built. The defenses are not ready. — https://nitter.net/alex_prompter/status/2040731938751914065#m
This didn't receive the attention it deserved. They pre-trained this model completely peer 2 peer, no data-centers. Everything was done over a permissionless network, I have tried the model, it's honestly not a good LLM but that's beyond the point. We NEED this, we NEED an alternative. – Download OpenCode – Download Pi – Pay for OpenSource – Share your AI sessions – Learn to do RL We can't be at the mercy of ANY lab. arxiv.org/abs/2603.08163
What if MLLMs could process visual data much faster without sacrificing performance? Eastern Institute of Technology, Ningbo, with USTC, SJTU, and LMU Munich presents HiDrop just for that! This new framework intelligently reduces visual tokens by processing them only when active fusion truly begins (Late Injection) and dynamically pruning them across deeper layers (Concave Pyramid Pruning with Early Exit). It focuses computation where it matters most. HiDrop compresses ~90% of visual tokens, matches original MLLM performance, and accelerates training by 1.72x. A new state-of-the-art for efficient MLLM training & inference! HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit Paper: arxiv.org/pdf/2602.23699 Code: github.com/EIT-NLP/HiDrop Our report: mp.weixin.qq.com/s/QKGZ7cFi0… 📬 #PapersAccepted by Jiqizhixin
There is a reason why NVIDIA is investing in brain/computer interfaces. Dustin (@r0ck3t23) Elon Musk just declared the human eye optional. Not improved. Not repaired. Not reconstructed. Optional. Musk: "Blindsight will enable those who have total loss of vision to be able to see again." That alone would be historic. Musk: "Including if they have lost their eyes, or the optic nerve." Eyes gone. Nerve gone. The entire optical pipeline physically missing from the skull. And the solution is not to rebuild what broke. It is to skip it entirely and wire synthetic signal straight into the visual cortex. Every surgery ever performed has tried to restore original hardware to factory condition. Neuralink does not restore. Neuralink treats the biological organ as optional infrastructure. Eye is gone. You do not rebuild the eye. You route around it. You stream raw visual data into the brain and let the cortex do what it was always doing anyway. Processing signal. Your eye never saw anything. Your brain saw. The eye was the middleman. It captured a narrow band of electromagnetic radiation and shipped it to the visual cortex. That is where the image was actually built. Neuralink is firing the middleman. Musk: "Maybe have never seen, were even blind from birth." A person who has never perceived a single photon of light. Given vision for the first time. Not through healing. Through hardware. And then Musk said the part that should rewire how you think about being human. Musk: "You can see in radar, you can see in infrared, ultraviolet." This is where it crosses from medical device to species upgrade. The human eye processes roughly 0.0035% of the electromagnetic spectrum. You are walking through [Translated from EN to English]
This article maps out some of the most important and influential papers on world model research from the past six months. nitter.net/robonaissance/status/2… Aviv Tamar (@AvivTamar1) Teaching a seminar on robot learning. Hit me with your favorite papers in the last 6 months (VLA, WM, RL, etc) — https://nitter.net/AvivTamar1/status/2041045100394905806#m
This is the Holodeck and why it is different than the Metaverse. Mark Zuckerberg’s metaverse can’t do this. Sadao Tokuyama (@tokufxug) 動画の1シーンレベルでの4DGS 動的シーン生成の新手法『TRiGS』 ガウシアンの出現・消失によるメモリ爆発を防ぐため剛体運動を統合。 ガウシアンの寿命を維持し無制限なメモリ増加を抑制します。 600〜1200フレームの長い動画でも高い忠実度と安定性を実現。 — https://nitter.net/tokufxug/status/2041061263778943035#m