In financial services, the ability to explain how a conclusion was reached matters as much as the conclusion itself. This agent uses LangSmith to preserve that decision log: every query issued, every response received, and every intermediate result produced before the final
LLMS
-
Rippling AI shipped millions using Deep Agents and LangSmith in 6 months
By
–
.
@Rippling AI runs on Deep Agents and LangSmith. Here’s how they shipped to millions of users in 6 months. -

Qwen3.7 Plus Released, Multimodal, Compared to GPT-5.4 and Opus 4.6
By
–
Qwen3.7 plus released. Looks good, but why do they compare their models to GPT-5.4 and Opus 4.6? Anyways, multimodal as well
-
Composer 2.5 now available in Grok Build, excels at complex tasks
By
–
Composer 2.5 is now available inside Grok Build.
— xAI (@xai) 1 juin 2026
Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions. pic.twitter.com/x7k4zVuVdWComposer 2.5 is now available inside Grok Build. Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions.
-

Reddit test: LLMs disagree on walk vs drive to car wash
By
–
Someone on Reddit asked 4 LLMs 100 times each: walk or drive to a car wash 100m away? Gemini 3.1 Pro said drive every time. Claude Opus 4.8 said walk every time.
-

GStack by Garry Tan goes viral, 100K stars on GitHub for dev cheat code
By
–
100K GITHUB STARS IN JUST A FEW WEEKS @GarryTan
’s GStack has gone completely viral, and for good reason. The YC CEO open-sourced his personal toolkit, and it's the ultimate cheat code for devs. It turns Claude Code from a basic chatbot into an entire virtual engineering -

Perplexity AI achieves new cost-performance frontier, outperforming Anthropic
By
–

It also sets a new cost-performance frontier. On DSQA it scores 0.871, ahead of Anthropic's 0.815, at nearly half the cost per task. On WideSearch it leads on score while running cheaper.
-

Search as Code replaces tool-calling with async search primitives
By
–
The traditional tool-calling approach suffers from high latency, manual control flow, and context pollution. With Search as Code, the model composes search primitives: fanning out queries asynchronously, deduping, filtering, joining, and ranking before results hits its context.
-

Self-Improving Language Models with Bidirectional Evolutionary Search
By
–
Most LLM search still works by sampling more rollouts or extending one path at a time. This paper's bidirectional evolutionary search does it in a smarter way. It breaks the task backward into smaller
-
AI agents rely on infrastructure; Ricoh runs Japanese models at 700+ tokens/sec
By
–
AI agents are only as useful as the infrastructure behind them. @ricoh is running custom Japanese AI models on SambaCloud at 700+ tokens/sec, delivering up to 10× faster performance than previous infrastructure. What once took a minute now finishes in ~10 seconds, making
