Graphics processing unit acceleration of the red/black SOR method. (23rd October 2012)
- Record Type:
- Journal Article
- Title:
- Graphics processing unit acceleration of the red/black SOR method. (23rd October 2012)
- Main Title:
- Graphics processing unit acceleration of the red/black SOR method
- Authors:
- Konstantinidis, Elias
Cotronis, Yiannis
Hidalgo, Jose Ignacio
Fernández‐de‐Vega, Francisco
Amor, Margarita
Doallo, Ramón
Fraguela, Basilio B.
Herrero, José R.
Quintana‐Ortí, Enrique S.
Strzodka, Robert - Abstract:
- <abstract abstract-type="main" id="cpe2952-abs-0001"> <title>SUMMARY</title> <p id="cpe2952-para-0001">This work presents our strategy, applied optimizations and results in our effort to exploit the computational capabilities of graphics processing units (GPUs) under the CUDA environment in order to solve the Laplacian PDE. The parallelizable red/black successive over‐relaxation (SOR) method was used. Additionally, a program for the CPU was developed as a performance reference. Various performance improvements were achieved by using optimization methods, which proved to provide significant speedup. Memory access patterns prove to be a critical factor in efficient program execution on GPUs and it is, therefore, appropriate to follow data reorganization to achieve the highest feasible memory throughput. The same approach exhibits performance benefits on the CPU version, as well. Eventually, a direct comparison of optimal versions' performance was realized. A 10 × speedup was measured for the CUDA version on an NVidia GTX480 GPU (NVidia Corp, Sta. Clara, CA, USA), exceeding 142 GB/s bandwidth, over the single threaded CPU version when run on an Intel Core i7 2600K CPU. The results prove that the global memory cache added on recent GPU architectures assist achieving high performance without requiring to employ the special memory types provided by the GPU (i.e. shared, texture or constant memory). Copyright © 2012 John Wiley & Sons, Ltd.</p> </abstract>
- Is Part Of:
- Concurrency and computation. Volume 25:Number 8(2013:Jun.)
- Journal:
- Concurrency and computation
- Issue:
- Volume 25:Number 8(2013:Jun.)
- Issue Display:
- Volume 25, Issue 8 (2013)
- Year:
- 2013
- Volume:
- 25
- Issue:
- 8
- Issue Sort Value:
- 2013-0025-0008-0000
- Page Start:
- 1107
- Page End:
- 1120
- Publication Date:
- 2012-10-23
- Subjects:
- Parallel processing (Electronic computers) -- Periodicals
Parallel computers -- Periodicals
004.35 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/cpe.2952 ↗
- Languages:
- English
- ISSNs:
- 1532-0626
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3405.622000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 4365.xml