Stock Research Agent V3 (Made by the LangChain Community) An AI platform transforming financial research, leveraging LangGraph and LangSmith for agent collaboration and real-time monitoring while delivering 73% cost savings. Check it out on GitHub!
LLMS
-

AI System Generates Historical Timelines Using LangGraph Multi-Agent Architecture
By
–
Event Deep Research (Made by the LangChain Community) An AI-powered system that generates historical timelines using LangGraph's multi-agent architecture. Automatically researches and compiles comprehensive JSON timelines of significant life events. Explore this
-
Snorkel AI Launches Terminal-Bench 2.0 and Harbor Tool
By
–
Congrats to the Terminal-Bench team from Snorkel AI! Thrilled to see Terminal-Bench 2.0, and excited to see Harbor — a game-changer. @laudeinstitute @stanfordailab @mike_a_merrill @alexgshaw @lschmidt3 @andykonwinsky @bradenjhancock
-

ERNIE-5.0-Preview-1022 Scores 1432 on LMArena
By
–


ERNIE-5.0-Preview-1022 from Baidu got a preliminary high ranking on LMArena and scored 1432 points. Feels like the gap is getting very small
-
Fine-tune DeepSeek-OCR Locally for Your Language
By
–
Fine-tune DeepSeek-OCR on your own language!
— Akshay 🚀 (@akshay_pachaar) 8 novembre 2025
(100% local)
DeepSeek-OCR is a 3B-parameter vision model that achieves 97% precision while using 10× fewer vision tokens than text-based LLMs.
It handles tables, papers, and handwriting without killing your GPU or budget.
Why it… pic.twitter.com/SfBXY4Is8oFine-tune DeepSeek-OCR on your own language! (100% local) DeepSeek-OCR is a 3B-parameter vision model that achieves 97% precision while using 10× fewer vision tokens than text-based LLMs. It handles tables, papers, and handwriting without killing your GPU or budget. Why it
-
Improved AI Responses with Visible Thinking Phase
By
–
I’m wondering if I accidentally flipped some feature flag or something (or maybe I’m in a beta test), because it definitely feels different. The results are way better, but having the “thinking” phase appear first makes it feel slower, even if the actual response time hasn’t
-

LLMs Maintain Consistent Mental Stability Across Preference Axioms
By
–
This figure visualizes the mental stability of models. LLMs like Qwen2.5 and Llama-3 stay consistent across all preference axioms transitivity, asymmetry, rating coherence proving they reason in structured, human-like ways. Even without training on user data, they infer what
-

LLM-as-a-Judge pipeline replaces traditional simulators with user history
By
–
This pipeline replaces traditional simulators with a simple idea: 1. Feed the model a short user history
2. Show two candidate slates
3. Ask: which one would this user prefer?
4. Aggregate responses across multiple LLMs That’s the “LLM-as-a-Judge” world model in action -

LLMs outperform random baselines in judging recommendation slates
By
–
This figure shows how well different LLMs judge slates (ordered recommendations) across datasets like Amazon, Spotify, MIND, and MovieLens. Lower “regret” = closer alignment with real user preferences. Turns out, LLMs consistently outperform random baselines when slates differ
-

LLMs judge your taste freakishly well in new recommender paper
By
–
Holy shit… LLMs just learned how to judge your taste and they’re freakishly good at it A new paper, “LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems,” just flipped recommender research on its head. Instead of simulating every click or dwell time, these
