← All articles
Published September 17, 2024by Felldude
A guide to training by the US Dept of Energy (For Real)
241 views8 reactions3 comments on CivitAI1 collected
amdflash-attentiontraining guidedeepspeedchat-gpt

Notice
Per US Public Law 91-651, Title 18 the Seal of the Department of energy has been altered so as not to mislead the public into thinking this article as being sponsored by the US government.
If the current president would like to consider a seat for the department of Lewdity I would be honored
The Article
Department of energy used AMD GPU's
Deepspeed Zero
Flash-Attention v2
My favorite excerpt:
"A Trillion parameter model requires a minimum of 14 Terabytes of memory, while an MI250X GPU has only 64 Gigabytes. So, to overcome the memory wall problem, we have explored a combination of model parallelism strategies."
Summary
the department of energy trained a 1 trillion parameter model (Chat GPT presumed size) with a fraction of the computing power thought needed.