They use "optimal" in a relative sense. I read it in a reply to my post, and then subsequently looked it up in LLM search engine—which happily told me why "Chinchilla is inference optimal because […]". But that's only relatively to undertrained models.
LLMS
-
Training AI Models with Extended 30k Context Windows
By
–
Worth mentioning that we can also train these models with very large context windows — up to 30k instead of 1k like the original.
-

Training from Scratch Pricing Table for Language Models
By
–
Here's a screenshot of our training-from-scratch pricing table that is on the webpage posted in the previous thread. We can also train to non-Chinchilla – do you have more information about the dataset?
-
Large Language Models Generate Cumulative Ambiguities Not Certainties
By
–
Humans have to generate a lot of words to try and explain the potential & challenges of large language #AI models https://
wsj.com/articles/chatg
pt-heralds-an-intellectual-revolution-enlightenment-artificial-intelligence-homo-technicus-technology-cognition-morality-philosophy-774331c6?mod=opinion_lead_pos5
…
"Enlightenment science accumulated certainties; the new AI generates cumulative ambiguities" -
LLMs Generating Training Data for Smaller Production Models
By
–
I was referring to NLP tasks that already exist pre LLM hype era where LLMs are typically used to few shot or zero shot generate a lot of data for smaller production models.
-
Production Performance vs Academic Benchmarks Eval Correlation
By
–
What production performance are you referring to then? Or are you referring to academic benchmarks being bad in general? Because you would always need some eval benchmark ideally correlated to the prod use case.
-
LLM Production Use Cases Beyond ChatGPT Benchmarking
By
–
It really depends on what you mean by production settings though. LLMs are used for many things other than being a ChatGPT. Overall I agree with your thread and benchmark results should be interpreted with caution. (But can be still useful some times)
-
Scaling Laws Limited Predictors of AI Model Capacity
By
–
yes, the scaling law, be it the chinchilla or Kaplan ones, are not at all a prediction/measure of the capacity of a model -in retrospect if you take a system-wide approach, with deployment/finetuning in consideration, they probably matter a lot less than people think
-
Compute Optimal Point Does Not Limit Performance Gains
By
–
I think compute optimal doesn't mean performance won't continue to increase if we continue to train a model even after the compute optimal point.
-

Open Source Chatbot Based on Hugging Face and Allen AI Models
By
–
Your personal Chatgpt based on open source models from @huggingface @allen_ai !