Yeah, this makes sense to me – it should be pretty easy to determine if something has RAG access or not by checking how it behaves on obscure reference-style questions
@simonw
-
GPT2-Chatbot RAG capabilities and model transparency concerns
By
–
Can anyone @lmsysorg confirm if gpt2-chatbot has the ability to run RAG against external tools or if it's working entirely from its own weights? A frustrating thing about opaque model releases is that without knowing details like this they're even harder to evaluate
-
Models Cannot Accurately Answer Questions About Themselves
By
–
I don't trust models to answer questions about themselves accurately
-
System Prompt Reliability vs Architecture Truthfulness in AI
By
–
It's interesting to me because I trust the system prompt to report a truthful cut-off date more than I trust it to provide a truthful architecture
-
Experiments on GPT2 Chatbot Context Length Limits
By
–
Anyone tried an experiment to see if they can figure out the context length for gpt2-chatbot?
-
New AI Model Development Timeline Testing Strategy
By
–
I would expect a completely new model to have an earlier cutoff date, because there's a bunch of extensive testing you have to do on that new model before you release it even as a preview Trained until Nov then 6 months of internal evaluation before public preview makes sense
-
New GPT Model Preview: Larger Parameters and Knowledge Base
By
–
From what I've seen so far this looks like a new model, trained differently from the current GPT-4 models – I think it has more "knowledge" baked in, so maybe a larger parameter count? I wouldn't be at all surprised if this turns out to be a preview of an OpenAI "GPT 4.5"
-
System Prompts Don’t Guarantee Truthful Model Self-Description
By
–
Worth noting that just because the system prompt says "based on the GPT-4 architecture" doesn't mean that the model is actually based on GPT-4! The goal of a system prompt is to influence the model to behave in certain ways, not to give it truthful information about itself
-
GPT2 Model Name Origin: OpenAI 2019 Release Joke
By
–
I'm confident the name is a joke – GPT2 was a model OpenAI released back in 2019
-
Benchmarking transparency needed for AI model testing
By
–
I'd prefer it if our benchmarking groups were 100% transparent about how their testing stack works and what models they are including in the mix