.
@Harvey
’s LAB benchmark approaches verification like a human would. Every task in a dataset has criteria for the task to pass. Legal agents can have 50+, with each one having its own judge call. It’s easy to audit, but can be expensive at scale. LangChain Labs teamed up with
AI
-

Harvey’s LAB benchmark uses human-like verification with per-task criteria
By
–
-

Capafy launches 5 e-commerce Skills from veteran operators
By
–

Capafy has released 5 pre-made e-commerce Skills, each built by an operator who has spent years on the store-side front line, with their hands-on playbook packaged into an agent that anyone can now run. The set covers > Commerce Ad Maker > Amazon Listing Image
-

Microsoft’s MAI-Image-2.5 takes #2 in Image Edit Arena
By
–
Microsoft just dropped MAI-Image-2.5 — and it immediately landed #2 in the Image Edit Arena (Single-Image-Edit) with a score of 1401. That's +10 pts over Nano Banana 2, Grok Imagine, and ChatGPT-Image-Latest-High Fidelity — and it pushes the Pareto frontier forward. Big W
-

SambaNova unveils disaggregated inference demo with 2x speedup
By
–
The first disaggregated inference demo for AI agents is now live. At #COMPUTEX2026, SambaNova demonstrated premium inference running in production at VC2 — using NVIDIA B200 GPUs for prefill and SambaNova RDUs for decode. The result: 2x faster inference than B200-only
-
AI Videos Quietly Automating Ad Industry with Claude and Airpost
By
–
The quiet revolution in AI videos isn’t Hollywood clips but something less fancy.
— AI Highlight (@AIHighlight) 3 juin 2026
It's the ad industry getting automated with tools like Claude and Airpost AI.
Let us explain: pic.twitter.com/fOF89rxKCDThe quiet revolution in AI videos isn’t Hollywood clips but something less fancy. It's the ad industry getting automated with tools like Claude and Airpost AI. Let us explain:
-

Microsoft MAI Guide: 1T Model, 35B Active, No Synthetic Data
By
–
Fantastic in depth guide about Microsoft MAI by @eliebakouch tl;dr about the model: Respect where respect is due. -zero synthetic data or distillation from previous models.
-1T model with 35B active, trained on 33.5T tokens -

8 LLMs for Agentic AI: Reasoning, Perception, Planning, Action
By
–
8 types of LLMs used in AI agents GPT • MoE • LRM • VLM • SLM • LAM • HRM • LCM Different models for reasoning, perception, planning, and action — not just chat. Agentic AI = model orchestration. #AI #LLMs #AgenticAI #GenAI #MachineLearning
-
@alphasignalai — 2026-06-03
By
–
YES, the CVE run makes that concrete, 100% accuracy at 85.1% fewer tokens, while the other systems stayed under 25%. Only word to push back on is "unprecedented" though, CodeAct was doing code-as-actions back at ICML 2024.
-
LangSmith Sandbox Gateway Observability
By
–
langsmith! Sandbox: https://
docs.langchain.com/langsmith/sand
boxes
… Gateway: https://
docs.langchain.com/langsmith/llm-
gateway
… Observability: https://
docs.langchain.com/langsmith/obse
rvability
… -
Uber limits coding agents to $1500/month per employee
By
–
Uber would now limit coding agents to $1,500/month per employee per tool – this seems reasonable to me, but it's also an interesting hint at the value Uber thinks these tools bring.