they didnt train new base model for 5 right what makes you so sure they will for 6 regardless my claim wasnt about new pretrains nor was it about timeline of 6 just enumerating what they clearly see as valuable for next steps to solve gdpval which is in my mind synonymous
LLMS
-
Hands-on LLM Engineering: Tokenization and Embeddings Projects
By
–
step-by-step LLM Engineering Projects each project = one concept learned the hard (i.e. real) way Tokenization & Embeddings > build byte-pair encoder + train your own subword vocab
> write a “token visualizer” to map words/chunks to IDs
> one-hot vs learned-embedding: plot -
OpenAI Training TFLOP Figures: Source and Transparency Questions
By
–
Where do those training TFLOP figure come from? I didn't know OpenAI shared numbers like that
-
OpenAI ChatGPT Pulse and Luma Ray3: Major AI Launches
By
–
Tomorrow’s AI News video is a big one for multiple reasons (including a fun announcement). Here are all the highlights from the past 2 weeks, what did I miss? – @OpenAI launched ChatGPT Pulse
– OpenAI unveiled age prediction + parental controls
– @LumaLabsAI released Ray3
– -

Flexible parsing unlocks hidden capabilities in weaker language models
By
–
Flexible parsing revealed hidden capabilities in weaker models LLM-based parsers captured valid answers from free-form reasoning, improving reported accuracy for models that struggled with rigid constraints. It’s a trade-off: exactness versus resilience.
-
Parsing and Evaluation: Technical Foundations for AI Systems
By
–
The bottom line: Parsing isn't just a technical detail. It's part of the evaluation story. Don’t forget to follow us @snorkelai and @realjustinbauer – and if you have questions about evals, RL environments, and expert-developed datasets, talk to us!
-

Structured Formats Limit AI Reasoning Capabilities
By
–
Structured formats can actually constrain reasoning Models like GPT-4.1 and Grok-3 performed worse when forced into rigid JSON structures. The formatting requirements limited their ability to think through complex problems.
-
Reasoning Models Show Resilience Across Parsing Methods
By
–
Reasoning-first models stayed resilient Claude Sonnet 4, Gemini 2.5 Pro, and o4-mini showed minimal sensitivity to parsing methods—their strong reasoning held steady across formats.
-
OpenAI Models Solve Programming Challenges With Code Sandbox
By
–
Confirmation from Ahmed El-Kishky – who worked on this project at OpenAI – that their models solved the programming challenges with a code execution sandbox but no internet access
-
Taking Extra Care with Model Decoding and Out-of-Distribution Constraints
By
–
okay maybe the question is – how would one take “extra care”? the D is the D, cant be touched. and yet people are constrained decoding thousands of tokens with presumably only slight intelligence drop. maybe my json example is too light of a constraint. the OODness would