Interesting launch! If the agent is good at "agentic search that doesn't stop until it finds what you need", consider evaluating on our BrowseComp benchmark, which measures just that! SimpleQA mainly targets models that don't browse: https://
openai.com/index/browseco
mp/
…
LLMS
-
BrowseComp Benchmark for Evaluating Agent Search Capabilities
By
–
-
Product Integration with LLM Technology
By
–
Nice exactly, good / clean summary. “Does your product speak LLM?”
-
Hybrid Architecture Combining Transformers and RNNs Advances
By
–
Wow this looks like a really important step. Combining the best bits of transformers and RNNs in one. Congrats!
-

AI fully generates a presentation of podcast lessons
By
–



Had an AI Agent watch a @hubermanlab podcast for me and compile all of the key takeaways into a presentation. These slides were 100% generated by AI.
-
Humorous suggestion about someone still using GPT-4o model
By
–
Man was still on 4o. We can start a gofundme for him
-
Man Uses ChatGPT for Relationship Advice During Lecture Hall Talk
By
–
Was attending a talk in a big lecture hall and the guy in front of me had the craziest conversation with ChatGPT for the whole hour about how to get his girlfriend back. Dozens of messages of pasting screenshots of text conversations to analyze tone of responses; whether to include an exclamation point to whether a smiley face was appropriately flirty. Apparently his ex-GF was talking to another guy in her lab and too stressed with work to give him attention, but also let him borrow her car, so he was getting mixed signals. Random stranger next to me was also spectating and found it so funny he literally cracked up in the middle of the talk. Convo ended with ChatGPT saying “you just had sex last week so youre no second class citizen.” Weirdest mix of pity, fascination, and awe I’ve ever felt
→ View original post on X — @_jasonwei, 2025-06-06 00:34 UTC
-
Model optimization tradeoffs and iterative improvements in AI development
By
–
building models is a juggling act, we optimized 05-06 for certain code tasks at the cost of other tasks, good lessons learned that the small regressions we saw in evals tracked to how people felt about the model, but we closed the gaps with 06-05!
-

Gemini now aware of FastHTML framework
By
–
Logan missed the BIG news here, which is that Gemini now knows about FastHTML! 😀
-
Cerebras Models on OpenRouter: 32K Context Expansion Coming
By
–
PSA: all Cerebras models are on OpenRouter with 32K context. We know you want longer context – team is on it!
-

Claude Gov: AI Models for US National Security
By
–
Introducing Claude Gov—a custom set of models built for U.S. national security customers. Already deployed by agencies at the highest level of U.S. national security, access to these models is limited to those who operate in classified environments.