I feel like we’re overdue for a revival here. Base models are still tedious to prompt, but now we have exactly the right tools for automating that work — aligned LLMs.
AI Dynamics
-

Reflections on LLM stylistic evolution
By
–
It’s crazy in hindsight that the stylistic mimicry of the GPT-2/3 era was such a local maximum for AI-written texts as cultural artifacts. For all their gains in making sense, aligned LLMs still can’t write with the eerie naturalism GPT-3 exhibited four years ago:
-
Questioning model language understanding versus ‘4o’
By
–
What kinds of language understanding is it struggling with? Also, this is relative to 4o?
-
Most models use instruct SFT and RL-based tuning
By
–
Most models that normal people have used in the past year (ChatGPT, Gemini, Claude, etc.) have some form of both instruct SFT and RL-based tuning. But yes I’m using “RLHF” inexactly in my post as a synecdoche for all post-training.
-
Pretraining and conditioning issues in instruct-tuned dialogue models
By
–
There’s conditioning from the dialog syntax that it’s being naively given in the same format that the instruct-tuned version receives. It’s seen these in pre-training, but the association isn’t strong enough apparently to make it act like a chatbot even most of the time.
-

Example contrasting base LLMs and RLHF-tuned models
By
–
I tried to write a prompt to show how base LLMs differ from the RLHF-tuned ones everyone knows, and I think this gives a bit of the flavor. A message from Llama 3.1 405B (base), on whether it’s useful to talk to base LLMs:
-

Discussion on RLHF vs post-training model comparisons
By
–
You’re not comparing RLHF to no RLHF here, you’re comparing different generations post-RLHF mystery models, likely of different sizes. If you don’t do post-training at all, naive attempts to talk to the model go like this:
-
On evaluating LLMs and token-scale confusions
By
–
Hard to say for LLMs because there’s no “all else being equal” — there aren’t any comparably large byte-token models. At smaller scales, I’d expect “Scunthorpe” style problems — confusions explainable as the result of spelling coincidences. I’m just guessing here though.
-
Identify Groqsters in Video Content
By
–
Spot the Groqsters in this video 👀 https://t.co/ZJo3TpXTEe
— Groq Inc (@GroqInc) 3 août 2024Spot the Groqsters in this video
-
Meta releases Llama 3.1 with new trust and safety tools
By
–
Build safe and responsible experiences with our latest Llama trust and safety tools. Along with Llama 3.1, we released Llama Guard 3, Prompt Guard and CyberSecEval 3. They’re available as part of the model download on our website and you can find more details here