Si tienes el hardware necesario a día de hoy ejecutar un LLM y tenerlo desplegado como una API no te lleva más de una tarde tonta.
AI HARDWARE
-
DataRobot Partners NVIDIA GPU-Accelerated Enterprise GenAI
By
–
Supercharge AI operations with DataRobot & NVIDIA! We partner with @nvidia to bring GPU-accelerated AI to the enterprise, enabling customers to build, govern, & operate #GenAI applications. Read their latest blog highlighting DataRobot as a key partner:
-

Infrastructure Setup and Scripts for 70B Model Deployment
By
–
From bare metal to a 70B model: infrastructure set-up and scripts https://
bit.ly/45XsNfi
#AI #MachineLearning #DeepLearning #LLMs #DataScience -

QLoRA Bug Consumes 4888GB VRAM in Llama 3 405B
By
–
One bug in bitsandbytes caused QLoRA to consume an extra *read notes* 4888 GB of VRAM. Even with QLoRA, fine-tuning Llama 3 405B on long sequence lengths ain't cheap.
-
Performance comparison of AI models on GPU vs LPU hardware
By
–
Users are also reporting that it is a bit slower than other models + they compare with Groq's performance (so the diff between GPU vs LPU is very visible)
-
Energy-Efficient AI Inference Acceleration Platform Recognized by Lazard
By
–
Honored to be named to @Lazard
's VGB AI Infra 40! Our energy-centric #AI #inference #acceleration addresses the critical need for efficient AI solutions. Proud to stand with fellow pioneers, driving sustainable, high-performance AI #computing forward. https://
lazard.com/research-insig
hts/venture-growth-banking-insights-lazard-vgb-ai-infra-40/
… -
Excitement for NVIDIA AI Foundry Capabilities and Innovations
By
–
We can't wait to see the incredible things enabled by NVIDIA AI Foundry!
-
NVIDIA AI Foundry Enables Custom Llama 3.1 Deployment
By
–
With NVIDIA AI Foundry, enterprises and nations can supercharge #generativeAI with a service that spans curation, synthetic data generation, fine-tuning, retrieval, guardrails, and evaluation to deploy custom @AIatMeta Llama 3.1 NVIDIA NIM Microservices. https://
nvda.ws/3WhKH7Q -
Llama 3.1 405B: Training the Largest Model at Scale
By
–
Training a model as large and capable as Llama 3.1 405B was no simple task. The model was trained on over 15 trillion tokens over the course of several months requiring over 16K @NVIDIA H100 GPUs — making it the first Llama model ever trained at this scale. We also used the 405B
-

Running 7B Parameter Models Locally on Apple Silicon
By
–
I am more and more excited about Local ML on Apple Silicon. Core ML in particular is starting to be a super nice stack to build on. Question: Do you want to have a 7B parameter model running at 30+ tokens/second, using less than 4GB of memory on your Mac? Then you need to