You can try deploying our models using our packaged solution (with vLLM), you'll get an API with control on temperature (which is how you get rid of repetition)
@arthurmensch
-
Mistral 7B Now Available in Production
By
–
Mistral 7B is now in prod, nice work @perplexity_ai ! https://t.co/mckgXbvPcw
— Arthur Mensch (@arthurmensch) 28 septembre 2023Mistral 7B is now in prod, nice work @perplexity_ai !
-
Mistral AI Releases First Model, Best Open Source 7B
By
–
At @MistralAI we're releasing our very first model, the best 7B in town (outperforming Llama 13B on all metrics, and good at code), Apache 2.0. We believe in open models and we'll push them to the frontier https://
mistral.ai/news/about-mis
tral-ai/
… Very proud of the team ! -
Llama II Release Advances Open-Source Language Model Progress
By
–
Great to see the release of Llama II, open-source LLMs are making good progress! Still a lot of room to improve OS models positioning on the efficiency/performance front — so that they eventually catch up with proprietary solutions. An interesting challenge
-
Mistral AI Founded: Guillaume Lample and Team Launch New Venture
By
–
Totally thrilled to be alongside @GuillaumeLample and @tlacroix6 to create Mistral AI. A lot of work ahead of us!
-
ChatGPT’s Ease with Boomer Requests Raises Questions
By
–
ChatGPT indeed seems at ease with boomer requests
-
Open vs Closed Source Operating Systems Competition Ahead
By
–
If that's indeed an OS, let's prepare to see an interesting replay of closed vs open-source operating systems in the coming years 🙂
-
Deep Learning vs Tree Methods on Small Datasets and Multitask Fine-tuning
By
–
I can see how deep learning methods may struggle to catch up with tree methods on small size datasets (and you say it in the thread). Wondering if you did try to do multitask fine-tuning on eg Transformers and saw a positive benefits? (we observe it in text, see eg Flan/T0)
-
Empirical Modelling: Approximation Theory and Optimization Analysis
By
–
Both are empirical modelling of experimental outcomes. The bottom one has some grounding in approximation theory and optimisation analysis. It also does not tend to 0 when N,D -> infinity, which is a sound property.