AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (117.9 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline

Fault-Tolerant Mechanism of the Distributed Cluster Computers

Yizi SHANG1Yang JIN2Baosheng WU1( )
State Key Laboratory of Hydroscience and Engineering, Tsinghua University, Beijing 100084, China
Department of Automation, Tsinghua University, Beijing 100084, China
Show Author Information

Abstract

The distributed system with high performance and stability is commonly adopted in large scale scientific and engineering computing. In this paper, we discuss a fault-tolerant mechanism under Linux circumstance to improve the fault-tolerant ability of the system, namely a scheme and frame to form the stable computing platform. In terms of the structure and function of the distributed system, active list and file invocation strategies are employed in the task management. System multilevel fault-tolerance can be achieved by repeated processes in a single node and task migration on multi-nodes. Manager node agent introduced in this paper administrates the nodes using the list, disposes of the tasks according to the nodes’ performance, and hence, to be able to make full use of the cluster resources. An evaluation method is proposed to appraise the performance. The analyzed results show the usefulness of the scheme proposed except for some additional overhead of memory consumption.

References

【1】
【1】
 
 
Tsinghua Science and Technology
Pages 186-191

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
SHANG Y, JIN Y, WU B. Fault-Tolerant Mechanism of the Distributed Cluster Computers. Tsinghua Science and Technology, 2007, 12(S1): 186-191. https://doi.org/10.1016/S1007-0214(07)70107-4

3

Views

0

Downloads

0

Crossref

N/A

Web of Science

0

Scopus

0

CSCD

Received: 01 February 2007
Published: 01 July 2007
© Tsinghua University Press 2007