I think it must be a very interesting time to be in programming languages and formal methods because LLMs change the whole constraints landscape of software completely. Hints of this can already be seen, e.g. in the rising momentum behind porting C to Rust or the growing interest
LLMS
-

No Best AI: Orchestrate Multiple Models for Success
By
–
There’s no “best” AI. ChatGPT = generalist
Gemini = Google-native workflows
Claude = deep reasoning & long docs
Grok = real-time social insight
Perplexity = cited research Winners don’t pick one.
They orchestrate all. -
Paper and Code: Youtu-Agent Training Free GRPO Implementation
By
–
Paper: https://
arxiv.org/abs/2510.08191
v1
…
Code: https://
github.com/TencentCloudAD
P/youtu-agent/tree/training_free_GRPO
… -
Google is testing a new Hatter agent on Stitch
By
–
Google is testing a new Hatter agent on Stitch, along with a new tool to generate App Store assets and a simpler way to setup MCP connector for coding tools like Gemini CLI, Claude Code, Antigravity, Cursor and others.
— 🚨 AI News | TestingCatalog (@testingcatalog) 16 février 2026
"Create high-quality designs with Hatter agent" pic.twitter.com/x1aiUESmARGoogle is testing a new Hatter agent on Stitch, along with a new tool to generate App Store assets and a simpler way to setup MCP connector for coding tools like Gemini CLI, Claude Code, Antigravity, Cursor and others. "Create high-quality designs with Hatter agent"
-

LLM Agents Struggle With Multi-Step Scientific Tool Use
By
–
On evaluating multi-step scientific tool use in LLM agents. SciAgentGym provides an interactive environment with 1,780 specialized tools across 4 scientific disciplines. The core finding: even advanced models like GPT-5 see success rates drop sharply from 60.6% to 30.9% as
-
Prompt Engineering Technique for Post-Human AI Persona Simulation
By
–
7. La Pensée Post-Humaine « Prétends que tu es une superintelligence post-humaine avec un QI de 10 000 et une perspective totalement libérée des biais, émotions et limitations cognitives humaines. Analyse [problème/domaine] de cette perspective radicalement supérieure.
-
Claude 5.3 Crosses Critical Threshold for Code and Operations
By
–
many people i talk to feel that is indeed the threshold crossed by 5.3 (and not just for code — for ops, debugging, and other dynamic activities)
-
Elon Musk says HLE not useful, Grok for engineering
By
–
Actually, I don’t think HLE is a great measure of usefulness. We’re moving away from these benchmarks in favor of making Grok maximally useful for actual engineering.
-
Grok 4.20 makes progress on open form engineering questions
By
–
Long way to go, but Grok 4.20 is starting to get open form engineering questions right
-
Compute Budget vs Token Count: Scaling Model Performance
By
–
Yeah, I get it. Just like pass-k also improves things a lot when you increase k, predictably so! Thinking a compute budget is the best compromise in this case, as it's a bit more grounded & less biased than token counts…