Accurate cross‒architecture performance modeling for sparse matrix‒vector multiplication (SpMV) on GPUs. (12th February 2014)
- Record Type:
- Journal Article
- Title:
- Accurate cross‒architecture performance modeling for sparse matrix‒vector multiplication (SpMV) on GPUs. (12th February 2014)
- Main Title:
- Accurate cross‒architecture performance modeling for sparse matrix‒vector multiplication (SpMV) on GPUs
- Authors:
- Guo, Ping
Wang, Liqiang
Limet, Sébastien
Smari, Waleed W.
Spalazzi, Luca
Hu, Jia
Gao, Jianliang - Abstract:
- <abstract abstract-type="main" id="cpe3217-abs-0001"> <title>Summary</title> <p id="cpe3217-para-0001">This paper presents an integrated analytical and profile‒based cross‒architecture performance modeling tool to specifically provide inter‒architecture performance prediction for Sparse Matrix‒Vector Multiplication (SpMV) on NVIDIA GPU architectures. To design and construct the tool, we investigate the inter‒architecture relative performance for multiple SpMV kernels. For a sparse matrix, based on its SpMV kernel performance measured on a reference architecture, our cross‒architecture performance modeling tool can accurately predict its SpMV kernel performance on a target architecture. The prediction results can effectively assist researchers in making choice of an appropriate architecture that best fits their needs from a wide range of available computing architectures. We evaluate our tool with 14 widely‒used sparse matrices on four GPU architectures: NVIDIA Tesla C2050, Tesla M2090, Tesla K20m, and GeForce GTX 295. In our experiments, Tesla C2050 works as the reference architecture, the other three are used as the target architectures. For Tesla M2090, the average performance differences between the predicted and measured SpMV kernel execution times for CSR, ELL, COO, and HYB SpMV kernels are 3.1<italic>%</italic>, 5.1<italic>%</italic>, 1.6<italic>%</italic>, and 5.6<italic>%</italic>, respectively. For Tesla K20m, they are 6.9<italic>%</italic>, 5.9<italic>%</italic>,<abstract abstract-type="main" id="cpe3217-abs-0001"> <title>Summary</title> <p id="cpe3217-para-0001">This paper presents an integrated analytical and profile‒based cross‒architecture performance modeling tool to specifically provide inter‒architecture performance prediction for Sparse Matrix‒Vector Multiplication (SpMV) on NVIDIA GPU architectures. To design and construct the tool, we investigate the inter‒architecture relative performance for multiple SpMV kernels. For a sparse matrix, based on its SpMV kernel performance measured on a reference architecture, our cross‒architecture performance modeling tool can accurately predict its SpMV kernel performance on a target architecture. The prediction results can effectively assist researchers in making choice of an appropriate architecture that best fits their needs from a wide range of available computing architectures. We evaluate our tool with 14 widely‒used sparse matrices on four GPU architectures: NVIDIA Tesla C2050, Tesla M2090, Tesla K20m, and GeForce GTX 295. In our experiments, Tesla C2050 works as the reference architecture, the other three are used as the target architectures. For Tesla M2090, the average performance differences between the predicted and measured SpMV kernel execution times for CSR, ELL, COO, and HYB SpMV kernels are 3.1<italic>%</italic>, 5.1<italic>%</italic>, 1.6<italic>%</italic>, and 5.6<italic>%</italic>, respectively. For Tesla K20m, they are 6.9<italic>%</italic>, 5.9<italic>%</italic>, 4.0<italic>%</italic>, and 6.6<italic>%</italic> on the average, respectively. For GeForce GTX 295, they are 5.9<italic>%</italic>, 5.8<italic>%</italic>, 3.8<italic>%</italic>, and 5.9<italic>%</italic> on the average, respectively. Copyright © 2014 John Wiley &amp; Sons, Ltd.</p> </abstract> … (more)
- Is Part Of:
- Concurrency and computation. Volume 27:Number 13(2015:Sep.)
- Journal:
- Concurrency and computation
- Issue:
- Volume 27:Number 13(2015:Sep.)
- Issue Display:
- Volume 27, Issue 13 (2015)
- Year:
- 2015
- Volume:
- 27
- Issue:
- 13
- Issue Sort Value:
- 2015-0027-0013-0000
- Page Start:
- 3281
- Page End:
- 3294
- Publication Date:
- 2014-02-12
- Subjects:
- Parallel processing (Electronic computers) -- Periodicals
Parallel computers -- Periodicals
004.35 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/cpe.3217 ↗
- Languages:
- English
- ISSNs:
- 1532-0626
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3405.622000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 4058.xml