EMERGENCY PAPER CLUB The @latentspacepod discord is meeting in 2hrs to talk thru @lvdmaaten et al's The Llama 3 Herd of Models, early contender to win the POTY* Awards! Join us (link below) with @swyx
, @vibhuuuus
, @picocreator
, @eugeneyan
, et al! *Paper of The Year, totally
LLMS
-

Emergency Paper Club Discusses Llama 3 Herd Models
By
–
-

Mistral Releases 123B Model: New Open-Weight Option Between 70B and 405B
By
–
And if you think there is too much distance between the 70B and the 405B the good folks at Mistral have now added a very nice 123B with some cool capabilities on code/math/multilingual What a week my friends, what a week for
open-weight models. And it’s only Wednesday -

Mistral Large 2 Now Available Free on Le Chat
By
–
You can use Mistral Large 2 on Le Chat — it's free! https://
chat.mistral.ai -

Mistral Large 2 Update Challenges GPT-4o and Sonnet
By
–
¡MISTRAL LARGE 2! Los franceses actualizan su modelo Large llevándolo a estar cerca de los mayores (GPT-4o, Sonnet 3.5). Es una mejora importante respecto al anterior modelo Mistral Large, pero claro… Muy a la par con Llama 3.1, que a diferencia de este modelo, es open
-

Mistral Large Improved Alignment and Instruction Capabilities Performance
By
–
Compared to the previous Mistral Large, much more effort was dedicated to alignment and instruction capabilities. On WildBench, ArenaHard, and MT Bench, it performs on par with the best models, while being significantly less verbose. (4/N)
-
Mistral Large Instruct 2407 Model Released on HuggingFace
By
–
The model is available (for research purposes only!) on HuggingFace: https://
huggingface.co/mistralai/Mist
ral-Large-Instruct-2407
…
Blog post: -

Mistral Large 2 outperforms Llama 3.1 on Multilingual MMLU
By
–
On Multilingual MMLU, the performance of Mistral Large 2 significantly outperforms Llama 3.1 70B base (+6.3% average over 9 languages) and is on par with Llama 3 405B (-0.4% below). (3/N)
-

Mistral Large 2 outperforms Llama 3.1 405B on coding benchmarks
By
–
On HumanEval and on MultiPL-E, Mistral Large 2 outperforms Llama 3.1 405B instruct, and scores just below GPT-4o. On MATH (0-shot, without CoT) it only falls behind GPT-4o.
(2/N) -
Mistral Large 2: New 123B Model Outperforms Llama 3.1 405B
By
–
Today, we release Mistral Large 2, the new version of our largest model. Mistral Large 2 is a 123B-parameter model with a 128k context window. On many benchmarks (notably in code generation and math), it is superior or on par with Llama 3.1 405B. Like Mistral NeMo, it was trained
-

Show Me The Prompt: Advanced LLM Prompt Engineering Guide
By
–
– Show Me The Prompt. https://
bit.ly/4cT6N7B
#AI #MachineLearning #DeepLearning #LLMs #DataScience