As AI research advances, more realistic software engineering benchmarks are critical to assess model performance and understand socioeconomic implications. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond.
AGENTS
-

SWE-Lancer: AI Models Tackle Engineering Tasks Across Full Stack
By
–
SWE-Lancer tasks span the full engineering stack, from UI/UX to systems design, and include a range of task types, from $50 bug fixes to $32,000 feature implementations. SWE-Lancer includes both independent engineering tasks and management tasks, where models choose between
-
Task Evals vs AGI: The Gap Between Simple Tests and Autonomous Agents
By
–
Let's keep in mind these are still super simple "task" evals. Little queries served on a platter, even if increasingly difficult. Which are super helpful, but when people talk about AGI they usually have an autonomous agent swarm in mind performing long-running jobs across
-
AI Systems Connecting and Reasoning Across Multiple Platforms
By
–
Love it. Totally inline with how we're thinking of the future as well on how AI can connect and reason across multiple systems.
-
BlackBox CyberCoder Transforms Enterprise Coding with SambaNova
By
–
◾️ @AiBlckbx is revolutionizing workflows with CyberCoder, their autonomous coding agent. With 10M+ users and Fortune 500 clients, they needed a fast and high-performance platform.
— SambaNova (@SambaNovaAI) 18 février 2025
🚀 SambaNova Cloud does exactly that.
Read the case study 👇#AI@blackboxai is revolutionizing workflows with CyberCoder, their autonomous coding agent. With 10M+ users and Fortune 500 clients, they needed a fast and high-performance platform. SambaNova Cloud does exactly that. Read the case study #AI
-
Function Descriptions from Docstrings in AI Tools
By
–
Credit also goes to Matthew Carrigan for the neat idea of getting function descriptions from docstrings: https://
huggingface.co/blog/unified-t
ool-use
… -

AISuite Simplifies LLM Function Calling for Developers
By
–
Announcing new aisuite capability: Easy function calling with LLMs! Function calling (tool use) is an important capability for agentic workflows and other LLM applications, but is cumbersome for developers to use (left column in image). Our open-source aisuite package simplifies
-

LangMem: Long-term Memory Integration for AI Applications
By
–
We've seen a lot of interest in long term memory and have spent a lot of time thinking about the best way to incorporate it into apps We've tried to distill some of these learnings into helper functions (and helper agents!) `pip install langmem` -> check it out
-
Deep Research Inventors Discuss AI Agents Innovation
By
–
pod: The Inventors of Deep Research! https://
latent.space/p/gdr While everyone cloning Deep Research, we asked @AarushSelvan and Mukund Sridhar, the original PM and Tech Lead who created the newest killer use case of AI Agents now copied by @openai
, @xai
, @perplexity_ai and a -

Grok3 vs ChatGPT: Is it the best AI? Comparative test
By
–
I just tested #Grok3 vs #ChatGPT and the results are rather… Surprising → https://youtu.be/jDDBVJDbCf0 Is it the best AI in the world as @ElonMusk says? We answer the question in this comparison.