Here's info on the estimated emissions for pre-training for the Gemma 2 family of models (from a paper published in July 2024). Those models are open-sourced, so should be easy for external parties to measure inference costs under various compute environments/settings (as
@jeffdean
-

Gemma 2 Pre-training Emissions Analysis Released
By
–
Here's info on the pre-training emissions of the Gemma 2 series of models, from a paper we released a few months ago (end of July 2024):
-

Gemma 2 Model Series: Emissions Estimates for AI Training
By
–
In July 2024, we released emissions estimates for pre-training the Gemma 2 model series (2B, 9B, and 27B parameter models, so ~12X, ~43X, and ~128X larger than the 213M parameter models you were looking at, and respectively trained on 2T, 8T, and 13T tokens of data). We estimate
-

Self-citation and attention mechanisms in academic papers
By
–
It sure seems to say that to an awful lot of people (all the "lifetime of five cars" stuff). Perhaps more importantly, in 2022, in https://
arxiv.org/pdf/2206.05229, a paper you co-authored, the self-citation description of your own 2019 work says "Attention was first drawn to the -

Google AI Model Training Power and Emissions Analysis Study
By
–
https://
arxiv.org/abs/2104.10350 that I helped co-author has quite a lot of details for models from Google and other models. It's actually quite a lot of work to do this. Table 4 in the paper has information about the training power, emissions, etc. for the Evolved Transformer NAS, as -
Geoff Hinton Wins Nobel Prize in Physics for Neural Networks
By
–
Congratulations to my good friend & former Google colleague Geoff Hinton for winning this year's Nobel Prize in Physics (along w/John Hopfield)! Geoff's impact on so many scientific fields continues to grow as neural networks are applied in more & more domains. Celebratory
-
Researcher Declines Co-authorship on AI Paper Correction
By
–
Indeed, and we invited @strubell to be a co-author with us on http://
arxiv.org/abs/2104.10350 that corrects the calculations done in the original paper, but she declined. -
Training Language Models Emissions Costs Misunderstood
By
–
It gives people a very incorrect view of the actual emissions costs of training language models. So people think "If a tiny model generated that much emissions, larger models must be thousands of times more than this" when in actual fact, because of the errors, this isn't the
-
Paper’s CO2 emissions claim for small transformer model corrected
By
–
Not sure what the more recent and much larger Gemma-2 models (2B, 9B and 27B parameters) have to do with this. The main issue is that your paper claimed that training a small 213m parameter transformer model emitted 284t CO2e, when in fact the correct number is 0.087 t CO2e, or
-
Emissions assessment errors in Evolved Transformer NAS research
By
–
I'm just looking at the timeline: 2019 – honest mistakes in Strubell et al. paper (
https://
arxiv.org/abs/1906.02243) in assessing emissions cost of Evolved Transformer neutral architecture search (off by 88X: 19X from not understanding the search process done in the actual So et al.