Parameter count and benchmark dominance aren't the same thing. Grok 4.20 proved that at 3 trillion parameters.
GENERATIVE AI
-
Copilot in Word: Retrofitting AI onto Legacy Software
By
–
Retrofitting Copilot onto Word is the software equivalent of adding electric motors to a horse carriage. The vehicle is still the wrong shape for what comes next.
-
AI Safety Concerns Block Broad Deployment Despite Strong Performance
By
–
93.9% SWE-bench and a 27-year-old OpenBSD bug found autonomously and the decision was still not to ship broadly. That's a data point about how seriously the internal assessment of the risks was taken.
-
Voxtral 4B Model: Advanced Text-to-Speech AI Release
By
–
Voxtral 4B Model: https://
huggingface.co/mistralai/Voxt
ral-4B-TTS-2603
… You can also try it here: https://
console.mistral.ai/build/audio/te
xt-to-speech/?utm_source=sumanth&utm_medium=x&utm_campaign=audio
… -

Mistral Open Sources Voxtral TTS Voice Cloning Model
By
–
Clone any voice with just a 3-second audio clip! Mistral just open sourced Voxtral TTS, a text-to-speech model that clones voices from 3 seconds of audio and runs on edge devices. Here's what makes it different. Most TTS models need cloud GPUs and long audio samples. Voxtral
-
Model Capacity and Data Memorization Risk in AI Systems
By
–
Pero a mayor capacidad del modelo más riesgo de memorización.
-
Volunteer needed for AI safety testing and vibe checks
By
–
i volunteer to vibe check all of the super dangerous models on normal tasks @AnthropicAI @OpenAI put me in the game!
-

Anthropic Model Card: Evaluation Overfitting Risks Assessment
By
–
2) Tal y como reportan en el propio Model Card hay riesgo de que estas evaluaciones hayan sido vistas por el modelo durante el pre-entrenamiento (a.k.a overfitting) y eso desvirtúa la interpretación de las métricas. Trabajo honesto el de Anthropic en la Model Card en muchos de
-

Power as Status Symbol: The AI Model Release Dilemma
By
–
the new status symbol is making a model so powerful you can’t release it
-

Mythos Model Evaluation: Why Single Benchmark Reporting Matters
By
–
Respecto a Mythos me han preguntado por qué en el vídeo de Youtube no he hecho mención a esta gráfica que todos estas comentando, y hay un par de motivos por el que descarté hablar de ello tras leer la Model Card. 1) Reportar la eficiencia de un modelo sobre un único benchmark