Our team's latest publication on Model Compression significantly reduces neural network size while maintaining state-of-the-art accuracy. Join us for the paper presentation at @ICCVConference on Friday, October 6th, at 10:30 am https://
go.axelera.ai/Differentiable
-Transportation-Pruning
… #ICCV2023
AI HARDWARE
-

Model Compression Reduces Neural Network Size Maintaining Accuracy
By
–
-
Nvidia CUDA Lock-In and Greedy Models Hardware Strategy
By
–
Covered some of this re: Nvidia’s (PyTorch) cuda lock-in, which is mostly for what I’d call “greedy models”—“low”-end hardware will be pretty good for a lot of inference (T4, A10, CPUs). Part of what makes Nvidia’s backing for PaxML interesting as well. https://
supervised.news/p/greedy-model
s-and-nvidias-open-source
… -
ChatGPT-Integrated Home Printer Transforms User Experience
By
–
The first home printer that integrates with ChatGPT is going to be wild No more buttons, weird flashing lights, etc. just send what you want printed and talk directly to the printer through chat.
-

LLaMa Inference Speedup 1.33x to 1.91x on GPUs
By
–
Speeding up LLaMa inference end-to-end by 1.33x on A6000 (for 13B model) and 1.91x on A100 (for 34b model). https://
arxiv.org/abs/2308.16369 -

Microsoft Research Builds Quantum Computer Transcending Binary Limitations
By
–
For a multidisciplinary group of researchers at the Microsoft Research Lab in Cambridge, U.K., the mission was to build a new kind of computer that would transcend the limitations of the binary systems in rapidly solving complex problems. https://
bit.ly/45L21FE -
Groq Challenges Nvidia’s AI Chip Dominance in LLM Inference
By
–
"Nvidia is the undisputed leader of AI chips… but as its competitors build out their own AI efforts, that’s all but guaranteed to decline… upstarts like Groq have shown advantages on certain models and workflows for LLM inference, today.” http://
groq.link/z -
DGX Station Hardware Failures and Support Experience
By
–
On the 4-GPU DGX Station, I had a general liquid cooling system failure, a CPU cooling failure, and a GPU failure. Support was great, and promptly replaced the system each time, but it still sucks.
-
GPU cluster reliability issues exceed conventional systems by orders of magnitude
By
–
I’m sure my thrice-replaced DGX station is an outlier, but the reliability reports I hear from big GPU cluster people are still grim. They have over an order of magnitude more failures than conventional systems. ML training is tolerant, but it is at a point that gets noticed.
-

Groq Day Event: Language Processing Unit Demo and LLM Performance
By
–
Join us at #GroqDay to learn more about the world’s first Language Processing Unit™, see a live demo of LLMs running on a Groq LPU™, and why fluid and fluent performance matters. Register at http://
groq.link/groqday5