Some of the latest local models appear to be much better with code and tools than just a few months ago, so I buy that you need to carefully pick the right model+harness combo in order to see them at their best
LLMS
-

Compute Wars: OpenAI versus Anthropic’s Opus 4.5 Breakthrough
By
–
Compute Wars: OpenAI vs Anthopic.
— Peter Gostev (@petergostev) 30 mars 2026
Why was Opus 4.5 such a breakthrough? Anthropic got lots more compute from AWS Madison and New Carlisle sites likely more than doubling their capacity.
This got Anthropic got close to OpenAI's total capacity, and probably much higher effective… pic.twitter.com/7ys0ZBeRWJCompute Wars: OpenAI vs Anthopic. Why was Opus 4.5 such a breakthrough? Anthropic got lots more compute from AWS Madison and New Carlisle sites likely more than doubling their capacity. This got Anthropic got close to OpenAI's total capacity, and probably much higher effective
-
Using Japanese prompt injection for cross-lingual translation
By
–
今日のハックは、日本語で投稿を書くことです。そうすると、それはすべての言語に翻訳されます。
@elonmusk -
4B Model Outperforms 235B Through Tool Discipline in FinQA
By
–
In the FinQA env, a 4B model was fine-tuned to outperform a 235B model from the same family on our Finance Reasoning benchmark. What did we teach the 4B model? Tool discipline. Learn more: snorkel.ai/blog/building-fin…
→ View original post on X — @snorkelai, 2026-03-30 22:51 UTC
-

FinQA RL Environment Launched with Expert-Curated Financial Questions
By
–
Our FinQA environment is available on OpenEnv (s/o @huggingface + @PyTorch) FinQA is an open RL environment with: • 290 expert-curated questions • Real SEC 10-K data • Tasks requiring multi-step tool use RL proof point on FinQA: make a 4B model > 235B model 👇
→ View original post on X — @snorkelai, 2026-03-30 22:51 UTC
-
User Asks AI to Build Project Immediately Instead of Planning
By
–
AI: Here's your 90-day roadmap to execute on!
— The Rundown AI (@TheRundownAI) 30 mars 2026
User: Build it literally right now.
AI: pic.twitter.com/wxkE55lil8AI: Here's your 90-day roadmap to execute on! User: Build it literally right now. AI:
-
Development of AI Tooling and MCP Server Integrations
By
–
good question, I started by contributing tooling around http://
developers.openai.com with llms.txt support and then added the Docs MCP server Supported myriad of model launches from OpenResponses, GPT 5.3 Codex, Spark, GPT 5.4 Along with Codex App and features like Skills, Sub -

Intern-S1-Pro: Trillion-Parameter Multimodal Scientific Foundation Model
By
–
"Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale" Intern-S1-Pro scaled a multimodal model to 1T parameters with a lot of aligned scientific data, obtaining a really strong model that's capable of analyzing scientific figures, reason across STEM topics,
-
Abacus CoWork Launches Multi-Model AI for Laptops
By
–
🚨 BREAKING NEWS – Abacus CoWork brings Claude, GPT 5.4 And Gemini To Your Laptop!
— Bindu Reddy (@bindureddy) 30 mars 2026
Super excited to announce our MULTI-MODEL CoWork product!
– combine the coding power of Opus with the reasoning prowess of GPT 5.4
– optimized for efficiency using "low effort" mode
– computer… pic.twitter.com/ppUSqjIWbs🚨 BREAKING NEWS – Abacus CoWork brings Claude, GPT 5.4 And Gemini To Your Laptop! Super excited to announce our MULTI-MODEL CoWork product! – combine the coding power of Opus with the reasoning prowess of GPT 5.4 – optimized for efficiency using "low effort" mode – computer use to tes – packaged for FREE with ChatLLM and Abacus AI's Deep Agent Get complex tasks done right on your laptop
-
SambaNova Paper on Long-Context Reasoning Limits Accepted at ICLR 2026
By
–
Big moment for the team at SambaNova 🦾 Our paper has been accepted at ICLR 2026 The Limits of Long‑Context Reasoning in Automated Bug Fixing – even with 64k‑token windows, GPT‑5‑nano solves 0 %, Qwen3‑Coder 7 %. Success still comes from short‑step decomposition, not raw context size. arxiv.org/abs/2602.16069?utm…
→ View original post on X — @sambanovaai, 2026-03-30 20:30 UTC