That was my argument to buy the DGX in the first place, but a committed cluster is a 6+ month deal, not an hourly purchase.
AI HARDWARE
-
Nvidia DGX Station GPU Failure and Form Factor Discontinuation
By
–
My $250k DGX station is being replaced for a second time after a GPU failed (first time was the cooling system). I now see that Nvidia has completely exited the form factor.
-
Regret over NVLink choice versus conventional PCI A100 setup
By
–
In hindsight, I would have been much better off with a more conventional system with PCI A100 cards — I still haven’t done anything dramatic with the extra NVLink bandwidth, and it could have had twice the GPUs.
-
GPT-2 Pre-training: Hardware Requirements and Token Processing Estimates
By
–
Rough example, a decent GPT-2 (124M) pre-training reproduction would be 1 node of 8x A100 40GB for 32 hours, processing 8 GPU * 16 batch size * 1024 block size * 500K iters = ~65B tokens. I suspect this wall clock can still be improved ~2-3X+ without getting too exotic.
-
GroqFlow: Automatic Toolflow for Machine Learning Workloads Mapping
By
–
GroqFlow™ is our automatic toolflow for mapping #machinelearning workloads to GroqChip™, built in support for #PyTorch, #Keras, #ONNX and #Hummingbird, with more coming in every update. Connect at http://
github.com/groq/groqflow/
discussions
… to request additional models or proof points. -

ON Semiconductor Faces Silicon Carbide Manufacturing Challenges
By
–
ON Semi struggles with silicon carbide, says William Blair https://
thetechnologyletter.com/the-posts/on-s
emi-struggles-with-silicon-carbide-says-william-blair
… // $ON $WOLF #investing #stocks #tech #siliconcarbide #semiconductors -
Cerebras Launches AI Model Studio for Transformer Training
By
–
ICYMI, Cerebras launched the AI Model Studio, a cloud-based pay-by-the-model computing service that makes training large transformer models fast, easy, and affordable. Read more in this Forbes article: https://
hubs.li/Q01xxQDx0 Sign up for a free trial: https://
hubs.li/Q01xxX3x0 -

Maxeler and Groq Accelerate Real-Time AI Solutions
By
–
Stop by our booth to hear how @MaxelerTech and @GroqInc can accelerate your systems with real time #AI solutions.
-
Convergence: Direct Computation and Machine Learning for Real-time AI
By
–
Convergence is combining direct computation with machine learning to deliver greater compute performance & lower latency for real-time systems.
Find out why convergence matters at @Business_AI
's #AISummit keynote Evolving Finance with Groq & Real-time AI: http://
spr.ly/60453KWiD -
Synthetic Gradients Inspire Tenstorrent Neural Network Research
By
–
#TenstorrentTop10: each week Tenstorrent will be highlighting a paper that has inspired our product development. This week's paper is Decoupled Neural Interfaces using Synthetic Gradients. #neuralnetwork #syntheticgradients # #research #tenstorrent https://
tenstorrent.com/research/tenst
orrent-top-10-decoupled-neural-interfaces-using-synthetic-gradients/
…