You're in a Research Scientist interview at OpenAI. The interviewer asks: "How would you expand the context length of an LLM from 2K to 128K tokens?" You: "I will fine-tune the model on longer docs with 128K context." Interview over. Here's what you missed:
RESEARCH
-

Harvey’s LAB benchmark uses human-like verification with per-task criteria
By
–
.
@Harvey
’s LAB benchmark approaches verification like a human would. Every task in a dataset has criteria for the task to pass. Legal agents can have 50+, with each one having its own judge call. It’s easy to audit, but can be expensive at scale. LangChain Labs teamed up with -

Microsoft’s MAI-Image-2.5 takes #2 in Image Edit Arena
By
–
Microsoft just dropped MAI-Image-2.5 — and it immediately landed #2 in the Image Edit Arena (Single-Image-Edit) with a score of 1401. That's +10 pts over Nano Banana 2, Grok Imagine, and ChatGPT-Image-Latest-High Fidelity — and it pushes the Pareto frontier forward. Big W
-

Microsoft MAI Guide: 1T Model, 35B Active, No Synthetic Data
By
–
Fantastic in depth guide about Microsoft MAI by @eliebakouch tl;dr about the model: Respect where respect is due. -zero synthetic data or distillation from previous models.
-1T model with 35B active, trained on 33.5T tokens -
@alphasignalai — 2026-06-03
By
–
YES, the CVE run makes that concrete, 100% accuracy at 85.1% fewer tokens, while the other systems stayed under 25%. Only word to push back on is "unprecedented" though, CodeAct was doing code-as-actions back at ICML 2024.
-

20 of 25 Top AI Researchers Say AI Will Soon Build AI
By
–
20 of 25 top AI researchers say AI will soon build AI. For decades, humans built every AI system from scratch. That assumption is quietly breaking down inside frontier labs. A new paper interviewed 25 top researchers from Stanford, OpenAI, Google DeepMind, and Anthropic.
-
SDPO explored in practical async setups praised
By
–
cool work folks – nice to see SDPO explored in practical/async setups!
-
GLM endpoint slow, mistakes help understanding, coding product fine
By
–
Yesterday evening GLM endpoint was slow, high traffic during U.S. working hours. Today it's making seemingly sillier mistakes, but every one it makes helps me understand the solution better — which has its benefits. Not perfect, but the subfrontier coding product is fine as-is.
-
Proceedings of Neuro for AI & AI for Neuro now available on PMLR
By
–
Volume 308 https://
proceedings.mlr.press/v308/ Proceedings of Neuro for AI & AI for Neuro Is now available on PMLR. -
BIGSET: Multi-agent AI for structured, verified datasets from web
By
–
Finding data from the web is easy. Formatting it is a nightmare.
— Charly Wargnier (@DataChaz) 3 juin 2026
🚨 Meet BIGSET by @Tiny_Fish.
A new multi-agent AI system that researches the web to build structured, verified datasets from one single prompt 🤯
Open source. BYOK support. Self-hostable.
4 wild use cases 🧵↓ pic.twitter.com/g4fiumZkASFinding data from the web is easy. Formatting it is a nightmare. Meet BIGSET by @Tiny_Fish
. A new multi-agent AI system that researches the web to build structured, verified datasets from one single prompt Open source. BYOK support. Self-hostable. 4 wild use cases ↓