Coordinated checkpoint versus message log for fault tolerant MPI. (6th February 2006)
- Record Type:
- Journal Article
- Title:
- Coordinated checkpoint versus message log for fault tolerant MPI. (6th February 2006)
- Main Title:
- Coordinated checkpoint versus message log for fault tolerant MPI
- Authors:
- Lemarinier, Pierre
Bouteiller, Aurelien
Krawezik, Geraud
Cappello, Franck - Abstract:
- Large clusters, high availability clusters and grid deployments often suffer from network, node or operating system faults and thus require the use of fault tolerant programming models. MPI is one of the most widely adopted programming models for high performance computing. There are several approaches for fault tolerance in an MPI environment. The automatic and transparent ones are based on either coordinated or uncoordinated checkpoint associated with a message log strategy. There are many protocols and optimisations for these approaches and several implementations have been made. However, few results of comparison between them exist. Coordinated checkpoint has the advantage of a very low overhead as long as the execution stays fault free. In contrast, uncoordinated checkpoint must be complemented by a message log protocol which adds a significant penalty for all message transfers even for fault free executions. The drawbacks of coordinated checkpoint are the synchronisation cost before the checkpoint, the synchronised checkpoint cost and the restart cost after a fault. Message log does not suffer from these problems, as it processes checkpoint and restart independently. These differences suggest that the best approach depends on the fault frequency. This paper investigates this question from a fair experimental protocol: we implement and test two protocols (coordinated checkpoint and pessimistic message log) on the same system and we compare them on a cluster according toLarge clusters, high availability clusters and grid deployments often suffer from network, node or operating system faults and thus require the use of fault tolerant programming models. MPI is one of the most widely adopted programming models for high performance computing. There are several approaches for fault tolerance in an MPI environment. The automatic and transparent ones are based on either coordinated or uncoordinated checkpoint associated with a message log strategy. There are many protocols and optimisations for these approaches and several implementations have been made. However, few results of comparison between them exist. Coordinated checkpoint has the advantage of a very low overhead as long as the execution stays fault free. In contrast, uncoordinated checkpoint must be complemented by a message log protocol which adds a significant penalty for all message transfers even for fault free executions. The drawbacks of coordinated checkpoint are the synchronisation cost before the checkpoint, the synchronised checkpoint cost and the restart cost after a fault. Message log does not suffer from these problems, as it processes checkpoint and restart independently. These differences suggest that the best approach depends on the fault frequency. This paper investigates this question from a fair experimental protocol: we implement and test two protocols (coordinated checkpoint and pessimistic message log) on the same system and we compare them on a cluster according to the frequency of faults that are generated artificially. The main conclusion is that uncoordinated checkpoint is relevant for a large scale cluster from one fault every hour for applications with large dataset. … (more)
- Is Part Of:
- International journal of high performance computing and networking. Volume 2:Number 2/3/4(2004)
- Journal:
- International journal of high performance computing and networking
- Issue:
- Volume 2:Number 2/3/4(2004)
- Issue Display:
- Volume 2, Issue 2/3/4 (2004)
- Year:
- 2004
- Volume:
- 2
- Issue:
- 2/3/4
- Issue Sort Value:
- 2004-0002-NaN-0000
- Page Start:
- 146
- Page End:
- 155
- Publication Date:
- 2006-02-06
- Subjects:
- fault tolerant MPI -- coordinated checkpoint -- message log -- high performance computing -- cluster computing
High performance computing -- Periodicals
Computer networks -- Periodicals
High performance computing
Periodicals
004.05 - Journal URLs:
- http://www.inderscience.com/jhome.php?jcode=ijhpcn ↗
http://www.metapress.com/openurl.asp?genre=journal&issn=1740-0562 ↗
http://www.inderscience.com/ ↗ - Languages:
- English
- ISSNs:
- 1740-0562
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 8687.xml