Guided installation of basic linear algebra routines in a cluster with manycore components. (3rd March 2017)
- Record Type:
- Journal Article
- Title:
- Guided installation of basic linear algebra routines in a cluster with manycore components. (3rd March 2017)
- Main Title:
- Guided installation of basic linear algebra routines in a cluster with manycore components
- Authors:
- Cuenca, J.
García, L. P.
Giménez, D.
Herrera, F. J. - Other Names:
- Lengauer Christian guestEditor.
Bougé Luc guestEditor.
Trystram Denis guestEditor.
Balaji Pavan guestEditor.
Leung Kai‐Cheung guestEditor. - Abstract:
- Summary: Computational systems are nowadays composed of basic computational components that share multiprocessors and coprocessors of different types, typically several graphics processing units (GPUs) or many integrated cores (MICs), and those computational components are combined in heterogeneous clusters of nodes with different characteristics, including coprocessors of different types, with varying numbers of nodes at different speeds. The software previously developed and optimized for simpler system needs to be redesigned and reoptimized for these new, more complex systems. The adaptation to hybrid multicore + multiGPU and multicore + multiMIC of autotuning techniques for basic linear algebra routines is analyzed. The matrix‐matrix multiplication kernel, which is optimized for different computational system components through guided experimentation, is studied. The routine is installed for each node in the cluster, and the information generated from individual installations may be used for a hierarchical installation in a cluster. The basic matrix‐matrix multiplication may, in turn, be used inside higher level routines, which delegate their efficient execution to the optimization of the lower level routine. Experimental results are satisfactory in different multicore + multiGPU and multicore + multiMIC systems. So the guided search of execution configurations for satisfactory execution times proves to be a useful tool for heterogeneous systems, where the complexity ofSummary: Computational systems are nowadays composed of basic computational components that share multiprocessors and coprocessors of different types, typically several graphics processing units (GPUs) or many integrated cores (MICs), and those computational components are combined in heterogeneous clusters of nodes with different characteristics, including coprocessors of different types, with varying numbers of nodes at different speeds. The software previously developed and optimized for simpler system needs to be redesigned and reoptimized for these new, more complex systems. The adaptation to hybrid multicore + multiGPU and multicore + multiMIC of autotuning techniques for basic linear algebra routines is analyzed. The matrix‐matrix multiplication kernel, which is optimized for different computational system components through guided experimentation, is studied. The routine is installed for each node in the cluster, and the information generated from individual installations may be used for a hierarchical installation in a cluster. The basic matrix‐matrix multiplication may, in turn, be used inside higher level routines, which delegate their efficient execution to the optimization of the lower level routine. Experimental results are satisfactory in different multicore + multiGPU and multicore + multiMIC systems. So the guided search of execution configurations for satisfactory execution times proves to be a useful tool for heterogeneous systems, where the complexity of the system means a correct use of highly efficient routines and libraries is difficult. … (more)
- Is Part Of:
- Concurrency and computation. Volume 29:Number 15(2017)
- Journal:
- Concurrency and computation
- Issue:
- Volume 29:Number 15(2017)
- Issue Display:
- Volume 29, Issue 15 (2017)
- Year:
- 2017
- Volume:
- 29
- Issue:
- 15
- Issue Sort Value:
- 2017-0029-0015-0000
- Page Start:
- n/a
- Page End:
- n/a
- Publication Date:
- 2017-03-03
- Subjects:
- autotuning -- heterogeneous computing -- hybrid programming -- parallel linear algebra -- manycore
Parallel processing (Electronic computers) -- Periodicals
Parallel computers -- Periodicals
004.35 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/cpe.4112 ↗
- Languages:
- English
- ISSNs:
- 1532-0626
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3405.622000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 2890.xml