MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU

(arxiv.org)

323 points | by chrsw 2 days ago ago

56 comments