Have you observed a meaningful difference between Q4 and Q2 either when it comes to tool calling? Would love to see how you measure that
LLMS
-

Cursor’s Composer 2 Competes with GPT-5.4 at 10-20x Lower Cost
By
–
Cursor's new in-house model now competes with GPT-5.4 and Opus 4.6 on coding. The difference: it's 10-20x cheaper to run. Composer 2 Fast output tokens: $7.50 / million. GPT-5.4 Fast: $75. Opus 4.6 Fast: $150. Terminal-Bench 2.0 scores: Composer 2 at 61.7, Opus 4.6 at 58.0,
-
LangSmith Fleet: Enterprise Workspace for Managing AI Agents
By
–
Introducing LangSmith Fleet: an enterprise workspace for creating, using, and managing your fleet of agents.
— LangChain (@LangChain) 19 mars 2026
Fleet agents have their own memory, access to a collection of tools and skills, and can be exposed through the communication channels your team uses every day.
Fleet… pic.twitter.com/bvrHqelj3eIntroducing LangSmith Fleet: an enterprise workspace for creating, using, and managing your fleet of agents. Fleet agents have their own memory, access to a collection of tools and skills, and can be exposed through the communication channels your team uses every day. Fleet
-
Frontier Models Rely on Memorization Over Generalizable Knowledge
By
–
This is more evidence that current frontier models remain completely reliant on content-level memorization, as opposed to higher-level generalizable knowledge (such as metalearning knowledge, problem-solving strategies…) https://t.co/QNqanOttqd
— François Chollet (@fchollet) 19 mars 2026This is more evidence that current frontier models remain completely reliant on content-level memorization, as opposed to higher-level generalizable knowledge (such as metalearning knowledge, problem-solving strategies…)
-
Open Source AI Inference: Competition and Community Contribution
By
–
as babyagi turns 3 yrs old, i finally sat down to compare the 9 iterations i did over the years…
— Yohei (@yoheinakajima) 19 mars 2026
this turned into https://t.co/x2evpDgcIM
a technical history of a personal project (which kind of captures the progress of the agent space overall) pic.twitter.com/L8fTiEAcM8as babyagi turns 3 yrs old, i finally sat down to compare the 9 iterations i did over the years… this turned into http://
babyagi.wiki a technical history of a personal project (which kind of captures the progress of the agent space overall) -
Cognitive Offloading Perils for LLM Users
By
–
*completely* wrong. the perils of cognitive offloading apply to all regular users of LLMs.
-

Continued Pretraining Boosts Model Quality and Cuts Serving Costs
By
–
We were able to significantly improve the model quality and cost to serve. These quality improvements come from our first continued pretraining run, providing a far stronger base to scale our reinforcement learning.
-

Frontier-Level Coding Model Pricing Tiers Revealed
By
–
It's frontier-level at coding, priced at: – Standard: $0.50/M input and $2.50/M output
– Fast: $1.50/M input and $7.50/M output -
Stay Ahead of AI by Mastering Frontier Models
By
–
How to never lose your job to AI: Just surf the models. Frontier models outclass humans at any form of knowledge that can be written down. But people who use frontier models in their field of expertise generate new, tacit, situational expertise that the models don't yet
-
LLM Performance Falls Apart on Unmemorizable Coding Benchmarks Due to Distribution Shift
By
–
Pretty shocking result (that once again confirms what I wrote about the perils of distribution shift, 25 years ago):
— Gary Marcus (@GaryMarcus) 19 mars 2026
Translate coding benchmarks into languages LLMs can’t memorize and performance utterly falls apart. https://t.co/wu5fh57nLZPretty shocking result (that once again confirms what I wrote about the perils of distribution shift, 25 years ago): Translate coding benchmarks into languages LLMs can’t memorize and performance utterly falls apart.