Read the paper here: https://
cdn.openai.com/papers/simpleq
a.pdf
…
Read the blog post here: https://
openai.com/index/introduc
ing-simpleqa/
…
Also see
@_jasonwei
-
OpenAI Introduces SimpleQA Paper and Blog Post
By
–
-

SimpleQA: New Open-Source Hallucinations Evaluation Benchmark
By
–
Excited to open-source a new hallucinations eval called SimpleQA! For a while it felt like there was no great benchmark for factuality, and so we created an eval that was simple, reliable, and easy-to-use for researchers. Main features of SimpleQA: 1. Very simple setup: there
-
OpenAI Researchers Share Strawberry Project Experience Video
By
–
22 minute video of OpenAI researchers talking about their experiences working on strawberry @woj_zaremba is very fun to work with @MillionInt is a legend
-
Meta-level Thinking in AI Research and RL Paradigm Shift
By
–
New talk from @hwchung27 about how to think "meta-level" in AI research. I have been impressed by Hyung Won's ability to identify new paradigms and totally give up any sunk cost. In late 2022 he realized the power of RL and has been preaching it ever since A fun story: when
-
Inverse Scaling and Model Performance on Different Prompt Types
By
–
1. I don't know of any great examples of inverse scaling (i.e., model performance gets much worse) off the top of my head, but I'm sure people will find some! You can see from our blog post that on some types of prompts like "personal writing", it seems like OpenAI o1-preview is
-
o1-mini Achieves Surprising 60% Score on AIME Math Competition
By
–
o1-mini is the most surprising research result i've seen in the past year obviously i cannot spill the secret, but a small model getting >60% on AIME math competition is so good that it's hard to believe congrats @ren_hongyu @shengjia_zhao for the great work!
-
OpenAI o1: Model That Thinks Before Answering
By
–
Super excited to finally share what I have been working on at OpenAI! o1 is a model that thinks before giving the final answer. In my own words, here are the biggest updates to the field of AI (see the blog post for more details): 1. Don’t do chain of thought purely via
-
OpenAI Engineer’s Competitive Advantage: Deep Code Understanding
By
–
Inspiring words from a young OpenAI engineer: “Why have I done well so far? I don’t think I’m smarter or more experienced than other people. But my competitive advantage is that I am willing to sit down and fully debug and completely understand code. I am willing to stay up late
-
Data Annotation Guidelines as Clear Communication Exercise
By
–
There is no better exercise in clear communication than writing data annotation guidelines
-
Soccer Team Roles Analogy for Tech Roles and Responsibilities
By
–
Similar analogies but in soccer:
– Goalie: Hardware folks
– Defender: Infra & data engineers
– Midfielder: All-around SWE and ML engineers
– Offense: General AI researchers
– Striker: Star AI researcher training the biggest models and doing yolo runs
– Coach: people managers