This time I'm tackling basics/intermediate concepts first. Not just the advanced stuff that many know me for. I've learned a lot in the past few years, having spent thousands of hours experimenting with prompting at this point.
GENERATIVE AI
-
Revisiting Prompting Material for New Updated Course and Workshops
By
–
BUT — I *am* revisiting a lot of that material right now though. That, and material from my advanced prompting course from 2023 and Build Powerful GPTs in 2024. Why? I've decided that putting out an updated prompting course with some live workshops is long overdue.
-
Pausing Thought Prompting Course Due to Uncertain Effectiveness
By
–
I still think thought prompting has potential, but I put the Thought Prompting course away without finishing it. I didn't want to teach people something that was just as likely to make their prompts worse as it was to make them better.
-
Three Complexity Bands: Where Reasoning Models Excel
By
–
where thinking models do well. Somewhere in the middle. See, there are three distinct complexity bands: → Low: Non-thinking LLMs actually outscore reasoning models.
→ Medium: The reasoning model's chain-of-thought helps.
→ High: Both vanilla and reasoning models drop to 0 %. -
Reasoning Models and the Goldilocks Zone of AI Thinking
By
–
Fast forward to today and we now have reasoning models ("LRMs") that are trained to think. When these first came out, I was skeptical. In my experience, LLM thinking was flawed. Apple's new paper confirms a similar finding. In fact, the identify a "Goldilocks band"…
-
Thought Prompts: Mixed Results in AI Output Quality
By
–
About half of the time, thought prompts worked well — better than without the thinking. And about half the time the thinking process actually made the output WORSE.
-
AI Brings Old Black-and-White Family Photos to Life in Color Video
By
–
J'ai testé sur des photo de famille. C'est bluffant…
Je l'ai ensuite passé à une IA pour en faire une vidéo… incroyable les vielles photos en noir et blanc prennent vie en vidéo en couleur -

RewardBench 2: New Multi-Skill Reward Model Evaluation Benchmark
By
–
9. RewardBench 2 RewardBench 2 is a new multi-skill benchmark for evaluating reward models with more challenging human prompts and stronger correlation to downstream performance.
-
LLM Memorization Capacity: Quantifying 3.6 Bits per Parameter
By
–
10. Memorization in LLMs This study introduces a method to quantify how much a model memorizes versus generalizes, estimating GPT models have a capacity of ~3.6 bits per parameter.
-
AlphaOne: Universal Framework for Controlling LRM Reasoning
By
–
7. AlphaOne Introduces a universal framework, α1, for modulating the reasoning progress of large reasoning models (LRMs) during inference. Rather than relying on rigid or automatic schedules, α1 explicitly controls when and how models engage in “slow thinking” using a tunable