gpt 5.5 and github pages and markdown
MULTIMODAL AI
-

Microsoft Research Paper on Agent-Based Interpretability for AI
By
–
NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets delegated to agents, the right target of interpretability shifts. This paper is a recipe for designing tools that agents can actually
-

Caltech AI Researchers Introduce Conversational Image Segmentation (CIS)
By
–
QUÉBEC ♡ AI / AI A new frontier is live.
A new frontier is opening. Frontier AI — http://
QUEBEC.AI:
models → governed intelligence.
models → governed intelligence. https://
quebecartificialintelligence.com/frontier-ai Frontier. AI‑First. Sovereign. #QuebecAI -

Ouster’s Native Color SPAD Pixels for Visible Light Capture
By
–
we need more benchmarks! awesome work by harvey here, and excited to work with them
-
ElevenAgents: Multi-Channel AI Agents for Customer Engagement
By
–
One agent, every channel, every modality. Meet your customers where they are with ElevenAgents. Learn more:
-

AI Agents Maintain Full Context Across Channels and Modalities
By
–
Agents maintain full context across channels and modalities. An agent on a voice call can send a message mid-conversation to confirm an appointment or share a document – then process what comes back, all in the same interaction.
-
ElevenAgents Introduce New Modalities for Handling Diverse Customer Inputs
By
–
Introducing new modalities for ElevenAgents
— ElevenLabs (@ElevenLabs) 6 mai 2026
Your customers don't just talk or type. They send photos, files, voice notes, and locations, and reach out across channels. Now your agents handle all of it. pic.twitter.com/dx4B4GchPuIntroducing new modalities for ElevenAgents Your customers don't just talk or type. They send photos, files, voice notes, and locations, and reach out across channels. Now your agents handle all of it.
-

ElevenAgents AI Processes Multimodal Data and Automates Customer Support
By
–
On WhatsApp, ElevenAgents now processes images, PDFs, audio messages, contacts, and location pins. A customer says "my wifi keeps dropping" and sends a photo of their router. The agent reads the indicator lights, runs a line check, and books an engineer visit without human
-

Uni-1 adds reasoning before image generation to resolve ambiguity.
By
–
Luma just released Uni-1, an image generation model that reasons first! The shift: image generation models typically work prompt-to-pixels. Uni-1 adds a reasoning step between input and output. It interprets creative direction, resolves ambiguity, then generates. That reasoning
-

Google déploie l’Intelligence Personnelle pour Gemini en UE
By
–

Google has started rolling out Personal Intelligence for Gemini in EU. Gemini Live will get it soon as well. > Gemini remembers your past chats so that you don't have to repeat yourself. Coming soon to Live.