Using structured weight pruning and knowledge distillation, the @NVIDIAAI research team refined Llama 3.1 8B into a new Llama-3.1-Minitron 4B. They're releasing the new models on @huggingface and shared a deep dive on how they did it https://
go.fb.me/b2h2c8
LLMS
-

NVIDIA Refines Llama 3.1 8B into Minitron 4B Model
By
–
-

Aya Model Paper: Cohere’s Latest AI Research
By
–
Learn more about the Aya Model paper: https://
cohere.com/research/paper
s/aya-model-paper-2024-02-13
… -

Grok 2 Reaches State-of-the-Art as AI Advances Across Multiple Domains
By
–
Top stories in AI today: -Grok 2 reaches state-of-the-art status
-Apple’s iPad is getting a robotic arm
-Creating sound effects with text
-Google’s Imagen 3 tops Midjourney, DALL-E
-5 new AI tools & 4 new AI jobs Read more: http://
therundown.ai/p/grok-2-reach
es-state-of-the-art-status
… -
Grok’s Future: RAG and Tweet Quality as Key Competitive Advantages
By
–
Mis conclusiones tras el vídeo de ayer sobre Grok, es que si xAI para la siguiente iteración se centra en mejorar el RAG/búsqueda de tweets mas que el LLM, podrán tener un producto bastante potente. Eso sí, también tendrán que mantener la calidad de los tweets de la plataforma.
-
Why “fine-tuning” is a misleading term for LLM updates
By
–
“Fine-tuning” is an unfortunate term. FT makes “fine” updates to model weights, but these can effect radical changes in almost any aspect of LLM behavior — to a user it’s more like “retraining.” This makes any use of FT’s lay definition (“tweaking”) in AI contexts confusing.
-
8B Model in 8-bit vs 4B Model in BF16 Precision
By
–
Yes it'd be interesting to see the 8B model in 8-bit precision vs. the 4B model in BF16.
-
Phi-2 vs Phi-3: Model Comparison and Benchmark Relevance
By
–
I think they just want to compare it to the 4B model. But I agree, phi-2 isn't the best comparison point. It's funny it's still regularly used in benchmarks despite the release of phi-3.
-
Models using cached tokens greatly reduce cost and latency
By
–
This feature really does change things. Gemini had cache, but only for costs — Claude claims 90% less cost *and* 85% less latency for cached tokens. Even if it can’t replace all SFT, it’s so much easier to iterate on you probably want to exhaust caching-based strategies first.
-

InfinityMATH: Scalable Dataset for Mathematical Reasoning
By
–
InfinityMATH A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning discussion: https://
huggingface.co/papers/2408.07
089
… Recent advancements in Chain-of-Thoughts (CoT) and Program-of-Thoughts (PoT) methods have greatly enhanced language models' mathematical reasoning -
Why is Haiku 97% cheaper than other AI models?
By
–
Why is haiku 97% cheaper? Is it not the same relative pricing across models?