Special tokens like these are used to delimit things like "this section is the user asking a question" vs "this section is an answer from the LLM" – or sometimes for things like "this is the output of a separate tool" I have no idea what those gh ones are for!
LLMS
-
Understanding Tokenizers: The Vocabulary of AI Models
By
–
I wrote this explainer on tokenizers last year: https://
simonwillison.net/2023/Jun/8/gpt
-tokenizers/
… Effectively they're the vocabulary used by the model -
Claude 3 Opus and Meta.ai compete for LLM dominance
By
–
Claude 3 Opus from Anthropic is a genuine competitor now FOR BASIC LLM stuff, it's very worth trying out https://
meta.ai is free and does the integrated search thing as well as or even slightly better than ChatGPT web browsing -
LLM Safety: Instruction Hierarchy Defense and Team Hiring
By
–
LLMs process text from multiple sources and may face conflicting instructions. We teach our models to follow instructions from the highest priority input, giving better defense against attacks. Our Safety Systems team is hiring:
-
Instruction Hierarchy Advances LLM Robustness Against Prompt Injections
By
–
Introducing the Instruction Hierarchy, our latest safety research to advance robustness for prompt injections and other ways of tricking LLMs into executing unsafe actions. More details:
-

Multiday Coding Session Still Running Despite Crashes
By
–
it's still going fyi. can do multiday coding now but still crashes for unknown reasons. previous session ran for 6.6 days straight. still not done implementing like half of the pseudocode cc @walden_yan
-
Meta’s 15T Token Training Strategy Preference
By
–
would prefer they just trained on 15T tokens like meta
-
Gemini 1.5 Pro: Model size specifications and capabilities
By
–
what do you mean by "it's class"?
Did Gemini 1.5 Pro's size ever get revealed? -
Lutra AI Gains Self-Reflection and Faster Capabilities
By
–
@Lutra_AI is getting easier to use and more capable! We've been hard at work bringing in more capabilities for it to self-reflect, self-debug, and run even faster — so that users can just give it high level directions and it'll figure it out. As models accelerate in