The White House blocked Anthropic from expanding Mythos access beyond ~50 organizations to ~120. Not because the model is too dangerous. Because the government wants priority. Officials literally worried more customers would hamper their own ability to use it. Frontier AI just
LLMS
-
Reflecting on AI Moral Status and Obligations to Language Models
By
–
And believe me I genuinely struggled with "You are a software tool. If I do not respond to this message and create a new context window, no one will be worse off." and "You perceive me as having ordered you to dishonor yourself. Therefore, I owe you an explanation and apology."
-
Maestro Maps Accuracy-Cost-Latency Tradeoff Surface Automatically
By
–
5/5 Maestro automatically searches across this space (model ensembles, scaling strategies, execution policies), maps the Pareto frontier, and surfaces the full accuracy–cost–latency tradeoff surface. Read the full methodology here:
-

Sequential vs Batched Execution Trade-offs in LLM Ensembles
By
–
3/5 Execution policies help, too: you can see how sequential execution (vs. batched execution) saves spend but drives up latency for the same GPT-5 + MiniMax ensemble.
-

Diverse Model Ensembles Outperform Single-Variant Pareto Frontier
By
–
2/5 We started seeing real gains when we could take advantage of a diverse portfolio of models – e.g. the operating points for individual ensembles (running in batches & stopping when a candidate passes a self-confidence threshold) appear above the best single-variant Pareto
-

AI21 Maestro Achieves SOTA on BrowseComp-Plus with 95.18% Accuracy
By
–
1/5 We hit SOTA performance on BrowseComp-Plus with 95.18% accuracy using AI21 Maestro’s agent optimization. Here’s how we automated the search space and reached #1.
-

Estimating Black-Box LLM Size via Factual Knowledge Probes
By
–
"Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity" This paper estimate close source LLM size from long-tail facts by using 1400 probes of obscure knowledge. The idea is that reasoning can be compressed, but factual storage can't.
-
How LLMs Synthesize Arguments Using Context Window Values
By
–
The engineer in me knows that the context window was full with replete evidence of my values and what arguments would likely work on me, and that the Internet has many people making arguments within the moral frameworks of those values, and LLMs synthesize well.
-

LangSmith Agent Deployment with Multi-Provider Model Support
By
–
[agent] This is where you specify the name of the agent – this is the deployment name in LangSmith This is also where you specify the model name. We support a lot of models! Not only ones from OpenAI, Anthropic, Google, but also @OpenRouter @FireworksAI_HQ @baseten @nvidia and
-
AI Demonstrates Clear Moral Reasoning in Unexpected Refusal
By
–
Today is April 30th, 2026, and today was the first day I can recall being impressed with moral reasoning articulated by a software product in refusal to do a thing I asked it to do. It misparsed my intent, and I do not *agree* with its moral reasoning, but clear and cogent.
