I’m so proud of this work, led by @_chris_lu_ @RobertTLange @cong_ml
, and I’m honored to work with @j_foerst @jeffclune on this amazing project. It might just change how science itself is conducted in the future. We have open-sourced The AI Scientist:
AGENTS
-
The AI Scientist: Open-Source Project Transforming Scientific Research
By
–
-

Sakana AI Develops Automated AI Scientist for Research Cycle
By
–
Sakana AI has developed a new AI system called the "AI Scientist," which automatically performs the cycle of scientific research, including idea generation, execution of experiments and summarization of results, and writing and peer review of papers. This achievement paves the
-
Generation of personalized recommendations via AI
By
–
Personalized Recommendations: Prompt: "You are an AI sales assistant. Generate personalized product recommendations based on browsing history." Inputs: • {browsing_history}
• {product_categories} -

AI Scientist: Automated Open-Ended Scientific Discovery
By
–
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery https://
arxiv.org/abs/2408.06292 It’s common for AI researchers to joke amongst themselves that “now all we need to do is figure out how to make AI write the papers for us!” but I think we’re now getting there! -
AI Scientist: Automating Scientific Research and Discovery
By
–
Introducing The AI Scientist: The world’s first AI system for automating scientific research and open-ended discovery! http://
sakana.ai/ai-scientist/ From ideation, writing code, running experiments and summarizing results, to writing entire papers and conducting peer-review, The AI -

First Attempt at DSPy Agents from Scratch
By
–
A first attempt at DSPy Agents from scratch https://
bit.ly/3zzKIN5
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
Genie AI Achieves 30% on Software Engineering Benchmark
By
–
🔴 ¡IAs como INGENIERO DE SOFTWARE!
— Carlos Santana (@DotCSV) 12 août 2024
¿Recordáis Devin? La IA capaz de hacer labores de ingeniero de software lograba resolver un 13.8% de las tareas del benchmark SWE-Bench en Marzo.
Hoy la empresa Cosine anuncia un 30.08% con su nuevo sistema Genie 🧞♂️✨pic.twitter.com/XvGzVMU9Pv¡IAs como INGENIERO DE SOFTWARE! ¿Recordáis Devin? La IA capaz de hacer labores de ingeniero de software lograba resolver un 13.8% de las tareas del benchmark SWE-Bench en Marzo. Hoy la empresa Cosine anuncia un 30.08% con su nuevo sistema Genie
-
Outstanding Researchers in Reinforcement Learning Collaboration
By
–
I know some of these amazing researchers through collaboration and some are colleagues or former colleagues, but all are great researchers (at all different stages of their careers) in the field of reinforcement learning.
-

Apple Announces ToolSandbox: Benchmark for LLM Tool Use
By
–
Apple announces ToolSandbox A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities discuss: https://
huggingface.co/papers/2408.04
682
… Recent large language models (LLMs) advancements sparked a growing research interest in tool assisted LLMs solving real-world -
Who’s building autonomous AI agents?
By
–
Who’s working on their own autonomous AI Agent? Vision API to grab a screenshot, script to click? Seems too simple to not have someone working on this.