A nice writeup about the TPUv4 system based on a talk my colleagues Norm Jouppi and Andy Swing delivered at @hotchipsorg 2023 on Tuesday. Learn more about the reconfigurable optical network!
AI HARDWARE
-

Groq LPU System Achieves 240 Tokens Per Second with Llama-2
By
–
Our LPU™ system is pushing the limits on LLM #inference perf again, now running Llama-2 70B at 240 tokens per sec per user! CEO @JonathanRoss321 shares more on the >2x improvement, why ultra-low latency matters, and if GPUs can still catch up. More at http://
groq.link/240tps -

Groq Platinum Sponsor at AI for Defense Summit 2023
By
–
We are proud to be a Platinum Sponsor at this year's AI for Defense Summit in National Harbor, MD. Stop by to see a demo of how Groq is enabling best-in-class performance with our software-defined deterministic chip, the Language Processing Unit™ (LPU™)
-

Generative AI at the Edge: Guest Lecture at Taiwan-Eindhoven
By
–
Join our very own Bram-Ernst Verhoef tomorrow for his guest lecture "Moving Generative AI Towards the Edge" at the inaugural Summer School Taiwan-Eindhoven, hosted at @TUeindhoven
. https://
etspss.nl/home -
Groq Day: AI Acceleration and LLM Inference Solutions
By
–
Join us at #GroqDay to understand why now is the moment that matters for AI acceleration of LLMs and how Groq inference solutions provide customer advantages. Register at http://
groq.link/groqday5 -
Image Models Limitations: Computational Constraints and Reduced Capability
By
–
Images are hundreds of thousands of pixels, so nobody can afford an architecture that runs a large model over every pixel. The resulting small models are in fact much stupider than GPT-4 and have trouble following even slightly complicated instructions.
-
Montreal Datacenter: Hardware Mix and CO2 Impact Timeline
By
–
I think Montreal is a relatively new datacenter (but great for CO2e!). The land purchase was just announced in May, 2021, so with construction and equipment deployment timelines, it may not be fully populated, so mix of hardware there may change.
-
Google Cloud Datacenters Carbon Intensity and TPUv4 Sustainability
By
–
Depends on the data center (I actually don't know which one this was filmed in). We report %age of Google carbon free energy as well as Grid carbon intensity
(gCO2eq/kWh) per Cloud datacenter here. Most TPUv4 pods are installed in ones marked "Low CO2" https://
cloud.google.com/sustainability
/region-carbon
… -

Groq showcases record-breaking LPU inference performance at ORNL conference
By
–
Stop by our table at @ORNL
's Smoky Mountains Computational Sciences & Engineering Conference to to get a sneak peak at the record breaking inference performance on the @GroqInc Language Processing Unit™ (LPU™) system, which we'll be announcing publicly tomorrow. -

Jais: Advanced Arabic LLM Trained on Condor Galaxy 1
By
–
(3/3) Jais was trained on the newly unveiled Condor Galaxy 1 (CG-1) AI supercomputer, built by the G42 – Cerebras strategic partnership. To learn more, check out the Press Release: https://
cerebras.net/press-release/
meet-jais-the-worlds-most-advanced-arabic-large-language-model-open-sourced-by-g42s-inception
…