DeepSeek-R1 proved RL boosts LLM reasoning. Meta now scales it for real-world SWE tasks Llama3-SWE-RL-70B, trained with RL on GitHub issues, PRs, & code fixes Outperforms SFT on math & code reasoning Best medium-sized LLM on SWE-Bench Verified
Comparable to GPT-4o
Meta’s Llama3-SWE-RL Outperforms GPT-4o on Code Tasks
By
–
