Running the "allegedly" leaked Mistral-Medium 70B locally. This is an extremely good model, which also confirms to be created by Mistral. If this is true, I feel bad for the guys at Mistral AI. It's also crazy how the M3 runs this at 10 tokens per second.
Mistral-Medium 70B Leaked Model Runs Locally at 10 Tokens/Second
By
–
