This week, MIT CSAIL will join other top ML researchers at ICLR to tackle a shift in focus from more powerful AI to more reliable systems Our papers at the conference show how to potentially make AI models stronger critical thinkers, more honest, & better at math
LLMS
-

MIT Harvard Study AI Agents Critical Thinking Battleship
By
–
“Collaborative Battleship” MIT & Harvard developed a collaborative version of Battleship to see if AI agents are as good at asking questions as answering them. They found that many LMs struggle w/critical thinking, but Monte Carlo inference strategies can help even tiny
-
Building Apps with AI APIs in Minutes at Low Cost
By
–
It takes less than 10 minutes to make an app out of X over on @Pokee_AI. Which includes the X API for free in its low token costs.
— Robert Scoble (@Scobleizer) 22 avril 2026
I talk to @AndrewWarner about why this is so powerful.
And it is cheaper now to do the same with your OpenClaw or Hermes or Manus or almost any AI… https://t.co/fc8Fnhf0AzIt takes less than 10 minutes to make an app out of X over on @Pokee_AI
. Which includes the X API for free in its low token costs. I talk to @AndrewWarner about why this is so powerful. And it is cheaper now to do the same with your OpenClaw or Hermes or Manus or almost any AI -

ChatGPT’s Limitations in Culinary Knowledge and Reasoning
By
–
ChatGPT doesn’t know its whisk from its elbow
-

Position Encoding: How Transformers Understand Data Order
By
–
Position Encoding: How Transformers Understand Order in Data
— Satya Mallick (@LearnOpenCV) 22 avril 2026
In this episode of Artificial Intelligence: Papers and Concepts, we explore Position Encoding, a fundamental concept that enables transformer models to understand the order of information. Since transformers process… pic.twitter.com/FS8RcTE7DtPosition Encoding: How Transformers Understand Order in Data In this episode of Artificial Intelligence: Papers and Concepts, we explore Position Encoding, a fundamental concept that enables transformer models to understand the order of information. Since transformers process
-
AI metrics reshape corporate performance reviews and hiring decisions
By
–
Been hearing wild stuff from folks inside big companies lately. Promotions, firings, and perf reviews are getting decided by tokens consumed and skills/MCPs connected. That’s the metric. That’s how they’re deciding who’s “good at AI.” It gets worse. People are literally running
-

Claude Code, Codex, Cursor: Comparing AI Model Sizes
By
–
Currently: – Claude Code is 4x bigger than Codex
– Codex is 2x bigger than Cursor
– Antigravity is almost as big as Cursor (but probably just because all Googlers use it? ) We'll add tagging for other agents asap. (Source: one data point from the @huggingface Hub team. Your -
NVIDIA and Google Expand AI Agent Solutions with Startups
By
–
NVIDIA Inception and Google for Startups are expanding with @coderabbitai and @FactoryAI using Nemotron‑based models on Google Cloud to power autonomous software agents, while @iamAible
, Mantis AI, @photoroom_ML and @baseten are building managed inference solutions powered by -
CrowdStrike Leverages NVIDIA NeMo for Cybersecurity AI Fine-tuning
By
–
.
@CrowdStrike uses NVIDIA NeMo open libraries on Gemini Enterprise Agent Platform to generate synthetic data and is fine-tuning Nemotron for domain-specific cybersecurity. -
ChatGPT Thinking vs API: How AI Models Process Information
By
–
It was Medium and I'm not 100% sure how thinking works in ChatGPT and whether it was equivalent – I would say it is equivalent to how the API works, so I think ChatGPT thinking would be different.