Mixtral 8x22B Instruct is out. It significantly outperforms existing open models, and only uses 39B active parameters (making it significantly faster than 70B models during inference). 1/n
@guillaumelample
-

Le Chat Temporarily Unavailable Due to High Request Volume
By
–
Due to an unexpected number of requests, Le Chat is temporarily unavailable. We apologize for the inconvenience — we are working on getting it back up and running as soon as we can, thanks for your patience!
-
Mistral Models Now Available on API with JSON and Function Calling
By
–
Our models are available on the @MistralAI API (La Plateforme). It supports JSON format and function calling. We are also making our commercial models available through Azure AI. Read more at: https://
mistral.ai/news/mistral-l
arge/
… https://
mistral.ai/news/le-chat-m
istral/
… Congrats to all the @MistralAI team -

Mistral Releases New Large Language Model with Enhanced Capabilities
By
–
Today, we are releasing Mistral Large, our latest model. Mistral Large is vastly superior to Mistral Medium, handles 32k tokens of context, and is natively fluent in English, French, Spanish, German, and Italian. We have also updated Mistral Small on our API to a model that is
-
Guillaume Lample praises Mistral AI team’s outstanding work and efficiency
By
–
Very proud of the small but amazing @MistralAI team for their outstanding work and building so quickly and efficiently. (n/n)
-

Mixtral Details and La plateforme Developer Platform Launch
By
–
More details about Mixtral can be found at https://
mistral.ai/news/mixtral-o
f-experts/
… We are also very happy to announce "La plateforme" our early developer platform (in beta & limited access), to access our models through our API: https://
mistral.ai/news/la-platef
orme/
… (7/n) -

Mixtral outperforms Mistral 7B in science, mathematics, and code
By
–
Compared to Mistral 7B, Mixtral is significantly stronger in science, in particular in mathematics and code generation. (5/n)
-

Mixtral Outperforms Llama 2 70B on European Language Benchmarks
By
–
Mixtral has been trained on a lot of multilingual data and significantly outperforms Llama 2 70B on French, German, Spanish, and Italian benchmarks. (4/n)
-
Mixtral: 12B Speed with 45B Parameter Access via Expert Selection
By
–
Even though each token only sees two experts, the selected experts can be different at each timestep. As a result, Mixtral decodes at the speed of a 12B model, while effectively having access to 45B parameters. (3/n)
-
Mixtral Architecture: 8 Feedforward Blocks with Router-Selected Experts
By
–
Mixtral has a similar architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks. For every token, at each layer, a router network selects two experts to process the current state and combine their outputs. (2/n)
