SWE-Lancer tasks more realistically capture the complexity of modern software engineering. Our tasks are full-stack and complex; the average task took freelancers over 21 days to resolve.
SOFTWARE
-
SWE-Lancer: New AI Coding Performance Benchmark Launched
By
–
Today we’re launching SWE-Lancer—a new, more realistic benchmark to evaluate the coding performance of AI models. SWE-Lancer includes over 1,400 freelance software engineering tasks from Upwork, valued at $1 million USD total in real-world payouts.
-
BlackBox CyberCoder Transforms Enterprise Coding with SambaNova
By
–
◾️ @AiBlckbx is revolutionizing workflows with CyberCoder, their autonomous coding agent. With 10M+ users and Fortune 500 clients, they needed a fast and high-performance platform.
— SambaNova (@SambaNovaAI) 18 février 2025
🚀 SambaNova Cloud does exactly that.
Read the case study 👇#AI@blackboxai is revolutionizing workflows with CyberCoder, their autonomous coding agent. With 10M+ users and Fortune 500 clients, they needed a fast and high-performance platform. SambaNova Cloud does exactly that. Read the case study #AI
-

LangMem: Long-term Memory Integration for AI Applications
By
–
We've seen a lot of interest in long term memory and have spent a lot of time thinking about the best way to incorporate it into apps We've tried to distill some of these learnings into helper functions (and helper agents!) `pip install langmem` -> check it out
-
Paylocity and Procore: Financial Software Investment Opportunities
By
–
Sifting financial software for gems: Paylocity, Procore Paylocity and Procore are two intriguing bets in the crowded market for HR, payroll and finance software, a group that has had very mixed results but that may see more M&A in the months and years to come.
-
Generative AI as Human Intelligence Co-pilot: AWS’s Fundamental Shift
By
–
2/ The buzzword of the year? Generative AI. But this isn’t hype—it’s a fundamental shift in how we work, create, and interact with technology. AWS is proving that #AI is more than just automation; it’s a co-pilot for human intelligence. #WAICF2025
-

Reproducing Software Build Issues: Testing Consistency Analysis
By
–
Hah yeah I can reproduce it and got something similar 5/5 attempts. I wonder what build they sent him earlier
-

Grok 3 crowned best LLM; XAI to widen lead, claims user.
By
–
For everyday users, the chatbot arena is the only benchmark that matters. Grok 3 is officially the best LLM. Given the speed at which @xai achieved this, they will only widen the gap over time.
-

xAI lance DeepSearch avec Grok 3
By
–
BREAKING : xAI is launching DeepSearch, the first AI Agent built on Grok 3
