Mixtral has a similar architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks. For every token, at each layer, a router network selects two experts to process the current state and combine their outputs. (2/n)
GENERATIVE AI
-
Mixtral: 12B Speed with 45B Parameter Access via Expert Selection
By
–
Even though each token only sees two experts, the selected experts can be different at each timestep. As a result, Mixtral decodes at the speed of a 12B model, while effectively having access to 45B parameters. (3/n)
-

Mixtral 8x7B: Open Weight Mixture of Experts Model Released
By
–
Very excited to release our second model, Mixtral 8x7B, an open weight mixture of experts model.
Mixtral matches or outperforms Llama 2 70B and GPT3.5 on most benchmarks, and has the inference speed of a 12B dense model. It supports a context length of 32k tokens. (1/n) -
New Research Initiative on General World Models Announced
By
–
Introducing General World Models.
— Runway (@runwayml) 11 décembre 2023
We believe the next major advancement in AI will come from systems that understand the visual world and its dynamics, which is why we’re starting a new long-term research effort around general world models.
Learn more: https://t.co/Z4JWm6dJjG pic.twitter.com/X8dW8fSv1YIntroducing General World Models. We believe the next major advancement in AI will come from systems that understand the visual world and its dynamics, which is why we’re starting a new long-term research effort around general world models. Learn more: http://
bit.ly/3RexmuJ -
Wife’s Favorite LLM: A Humorous Marriage Counseling Moment
By
–
My wife and I were at a married couples counseling meeting. The speaker was asking the audience random questions. Then he called on me and said, "Sir do you know what your wife's favorite LLM is?" Feeling kind of singled out, I turned to my wife and said, "It's Llama, isn't
-

WSJ Reviews The Coming Wave: AI Wonders and Troubles Ahead
By
–
review of http://
the-coming-wave.com from @WSJ https://
wsj.com/arts-culture/b
ooks/the-coming-wave-review-wonders-ahead-trouble-too-27e2fb18?mod=arts-culture_lead_story
… -
SDXL Fine-tune Model Based on Upscaled Emoji Data
By
–
We need an SDXL fine tune based on upscaled emoji
-
Leonardo vs MidJourney: Which generative AI tool to choose?
By
–
Why Leonardo over MJ? I haven’t used it in a long while, should check it out again
-
MoE Inference Benefits: Why Mixture of Experts Improves Performance
By
–
It can be unintuitive why the Transformer-style MoE (in Mixtral/GPT4) has inference benefits.
Dima simplifies it with a clear explanation showcasing that MoE help inference once there's sufficient volume of requests (which hopefully are diverse enough that they don't hit the same -
Welcome to the World Thomas: Born in AI’s Most Interesting Era
By
–
Congrats! Welcome to the world Thomas. You’re being born in the most interesting time in human history. Enjoy.