I'm very curious about OpenAI's planned intern researcher release by September this year. Having tried using current LLMs for OpenAI's Golf Challenge, I would say that Codex & Opus are actively bad researchers (not meaningful difference between them)
– They come up with small
OpenAI intern researcher release evaluation LLM research capabilities
By
–