Oh and btw big fan of @nvidia Nemotron! What makes NVIDIA Nemotron special is that it's not just another open-weight model. With Nemotron, NVIDIA provides not only the weights but also the training data, recipes, and technical details.
RESEARCH
-

OpenAI introduces LifeSciBench benchmark for life science research
By
–
Introducing LifeSciBench, a benchmark for measuring and improving how well AI supports real-world life science research. Developed with 173 scientists from biotechnology and pharmaceutical research, LifeSciBench includes 750 expert-authored tasks across seven biological research
-
LifeSciBench: A Foundation for Realistic AI Evaluation in Life Sciences
By
–
LifeSciBench is a foundation for more realistic evaluation, targeted improvements, and continued partnership with the life sciences community—helping the field measure progress, identify gaps, and improve AI together for the benefit of everyone.
-

LifeSciBench tests reasoning; GPT‑Rosalind outperforms GPT‑5.5
By
–
Benchmarks often test biological knowledge or narrow skills. The tasks in LifeSciBench test whether models can reason from evidence, work with scientific artifacts, handle uncertainty, and make useful decisions under real-world constraints. GPT‑Rosalind scores above GPT‑5.5
-
Vast knowledge and low intelligence are norm for AIs
By
–
Vast knowledge and low intelligence seldom coexist in a human, but they're they're the norm among AIs.
-
Improving a challenging reaction in medicinal chemistry with GPT-5.4
By
–
GPT-5.4 for improving a challenging reaction in medicinal chemistry: https://t.co/VqnADd8TP3
— Greg Brockman (@gdb) 17 juin 2026—
GPT-5.4 for improving a challenging reaction in medicinal chemistry:
— -

Harbor: standard framework for stateful agent evaluations
By
–
Harbor is an excellent framework for running longer, more stateful agent evaluations. It underpins Terminal Bench 2 and is becoming the industry standard. LangSmith sandboxes now integrate Harbor!
-
Grounded Reasoning Cup live at DataAISummit 2026
By
–
The Grounded Reasoning Cup is live at #DataAISummit 2026! Over the next few hours, 12 teams from leading universities go head-to-head in a live AI agent championship, tackling real-world enterprise reasoning challenges using models from @GoogleDeepMind
, @AnthropicAI
, and -
Economic exploitation of unique signals: LangChain Labs study
By
–
We partnered with @FireworksAI_HQ to answer the following question… How can we economically exploit important signals from each unique trace while maintaining state-of-the-art performance? Read our LangChain Labs study
-

Fine-tuning open models can surpass state-of-the-art models
By
–
Fine-tuning open models can surpass or match state-of-the-art models. Base
@Alibaba_Qwen
out of the box with good prompting:
Solid for perceived error classification, below state-of-the-art model performance. With a LoRA SFT job:
Both