Such a great set of launches at today’s DevDay; will unlock new kinds of applications! Excited to see what developers will build, and congrats to the OpenAI team for shipping.
LLMS
-

Meta’s Llama 3.2 Achieves Fastest Token Processing Speeds
By
–
Shout out to @ArtificialAnlys for independently verifying that we're the fastest on: Llama 3.2 1B with 2470 tokens per sec Llama 3.2 3B with 1566 tokens per sec @AIatMeta
-

SambaNova Cloud Achieves Fastest Llama Inference Speeds
By
–
Yes, we’re fast. In fact, the fastest! 🚀🚀
— SambaNova (@SambaNovaAI) 1 octobre 2024
SambaNova Cloud delivers the fastest inference on @AIatMeta's Llama 3.2 1B and 3B — all running at full-precision.
✅ 2470 tokens per sec on 1B
✅ 1566 tokens per sec on 3B#LLM #AI Start developing ⤵️Yes, we’re fast. In fact, the fastest! SambaNova Cloud delivers the fastest inference on @AIatMeta
's Llama 3.2 1B and 3B — all running at full-precision. 2470 tokens per sec on 1B 1566 tokens per sec on 3B #LLM #AI Start developing -

NVIDIA 72B Model Rivals Llama’s 405B Performance
By
–
Wow. New NVIDIA 72B model rivals Llama's 405B! https://
nvlm-project.github.io -
Model Distillation: Compressing Larger AI Models Efficiently
By
–
Model "Distillation": Compress larger models to smaller versions Cool. Do you know more about it?
-

Llama Stack Official Distribution: Unified API for Multiple Providers
By
–
We’re excited to share the first official distribution of Llama Stack! It packages multiple API Providers into a single endpoint for developers to enable a simple, consistent experience to work with Llama models across a range of deployments. Details https://
go.fb.me/xfi7g3 -

Comprehensive AI Taxonomy and Classification Categories
By
–
Cela pourrait aider pas mal de personnes :
-
OpenAI Releases Realtime API, Vision Fine-tuning, Prompt Caching
By
–
realtime api (speech-to-speech): https://
openai.com/index/introduc
ing-the-realtime-api/
… vision in the fine-tuning api: https://
openai.com/index/introduc
ing-vision-to-the-fine-tuning-api/
… prompt caching (50% discounts and faster processing for recently-seen input tokens): https://
openai.com/index/api-prom
pt-caching/
… model distillation (!!): https://
openai.com/index/api-mode
l-distillation/
… -
OpenAI real-time speech API threatens 20 million call center jobs
By
–
OpenAI just announced a real-time speech API (advanced voice mode) at Dev Day. The “speech wrappers” are coming to replace 20 million global call center jobs…
-

Lightning Fast AI Hackathon with $10k Prize Pool
By
–
⚡️ Introducing our Lightning Fast AI Hackathon ⚡️
— SambaNova (@SambaNovaAI) 1 octobre 2024
💰$10k Prize Pool!
Get ready to unleash your creativity with the incredible speed and versatility of the fastest #AI inference on the best open source models in the world! 🏎️💨
Think you can hack it? Learn more ⤵️Introducing our Lightning Fast AI Hackathon $10k Prize Pool! Get ready to unleash your creativity with the incredible speed and versatility of the fastest #AI inference on the best open source models in the world! Think you can hack it? Learn more