Efficient MPI‐AllReduce for large‐scale deep learning on GPU‐clusters. (9th December 2019)
- Record Type:
- Journal Article
- Title:
- Efficient MPI‐AllReduce for large‐scale deep learning on GPU‐clusters. (9th December 2019)
- Main Title:
- Efficient MPI‐AllReduce for large‐scale deep learning on GPU‐clusters
- Authors:
- Thao Nguyen, Truong
Wahib, Mohamed
Takano, Ryousei - Abstract:
- Summary: Training models on large‐scale GPUs‐accelerated clusters are becoming a commonplace due to the increase in complexity and size in deep learning models. One of the main challenges for distributed training is the collective communication overhead for large message sizes: up to hundreds of MB. In this paper, we propose two hierarchical distributed memory multileader AllReduce algorithms optimized for GPU‐accelerated clusters (named lr_lr and lr_rab ), in which GPUs inside a computing node perform an intra‐node communication phase to gather and store results of local reduced values to designated GPUs (known as node leaders). Node leaders then keep a role as an inter‐node communicator. Each leader exchanges one part of reduced values to the leaders of the other nodes in parallel. Hence, we are capable of significantly reducing the time for injecting data into the inter‐node network. We also overlap the inter‐node and intra‐node communication by implementing our proposal in a pipelined manner. We evaluate those algorithms on the discrete‐event simulation Simgrid. We show that our algorithms, lr_lr and lr_rab, can cut down the execution time of an AllReduce microbenchmark that uses the logical ring algorithm ( lr ) by up to 45% and 51%, respectively. With the pipelined implementation, our lr_lr_pipe achieves 15% performance improvement when compared with lr_lr . In addition, the simulation result also projects power savings for the network devices of up to 23% and 32%.
- Is Part Of:
- Concurrency and computation. Volume 33:Number 12(2021)
- Journal:
- Concurrency and computation
- Issue:
- Volume 33:Number 12(2021)
- Issue Display:
- Volume 33, Issue 12 (2021)
- Year:
- 2021
- Volume:
- 33
- Issue:
- 12
- Issue Sort Value:
- 2021-0033-0012-0000
- Page Start:
- n/a
- Page End:
- n/a
- Publication Date:
- 2019-12-09
- Subjects:
- distributed deep learning -- high‐performance computing (HPC) -- MPI -- AllReduce
Parallel processing (Electronic computers) -- Periodicals
Parallel computers -- Periodicals
004.35 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/cpe.5574 ↗
- Languages:
- English
- ISSNs:
- 1532-0626
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3405.622000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 17819.xml