I don't trust models to answer questions about themselves accurately
LLMS
-
Experiments on GPT2 Chatbot Context Length Limits
By
–
Anyone tried an experiment to see if they can figure out the context length for gpt2-chatbot?
-
New AI Model Development Timeline Testing Strategy
By
–
I would expect a completely new model to have an earlier cutoff date, because there's a bunch of extensive testing you have to do on that new model before you release it even as a preview Trained until Nov then 6 months of internal evaluation before public preview makes sense
-
New GPT Model Preview: Larger Parameters and Knowledge Base
By
–
From what I've seen so far this looks like a new model, trained differently from the current GPT-4 models – I think it has more "knowledge" baked in, so maybe a larger parameter count? I wouldn't be at all surprised if this turns out to be a preview of an OpenAI "GPT 4.5"
-
GPT2 Model Name Origin: OpenAI 2019 Release Joke
By
–
I'm confident the name is a joke – GPT2 was a model OpenAI released back in 2019
-
Predibase Launches 10x Faster Fine-tuning Stack with Llama3 Support
By
–
Announcement: New Improved #Finetuning Stack—10x Faster Training! Here are the highlights: New #training stack up to 10x faster + better model quality #Llama3 available for inference + fine-tuning New Python #SDK: More consistent and robust
-
GPT-4.5 disappoints: gpt2-chatbot outperforms expectations
By
–
gpt2-chatbot is good. really good. but if this is gpt-4.5, I’m disappointed.
-

AWS Cohere Command R deployment webinar for enterprise automation
By
–
AWS and Cohere will share more about Command R and R+ and how to seamlessly deploy our models at scale to automate critical business tasks. Join us on May 1st at 1:00 pm ET: https://
pages.awscloud.com/awsmp-gim-xlwy
-webinar-aim-generative-ai-cohere-spotlight-series.html?trk=475ad537-df52-4a80-9317-6ed4db81757a&sc_channel=el
… -
Command R Models Now Available on Amazon Bedrock
By
–
Our latest Command R model family is now available on Amazon Bedrock by @awscloud
! These state-of-the-art LLMs are designed to tackle enterprise-grade workloads, and excel at business-critical capabilities like multilingual coverage, RAG, and tool use.
