Application fault tolerance with armor middleware : Recovery-Oriented Computing

Zbigniew Kalbarczyk,Ravishankar K. Lyer,Long Wang
IF: 2.68
2005-01-01
IEEE Internet Computing
Abstract:Many current approaches to software-implemented fault tolerance (SIFT) rely on process replication, which is often prohibitively expensive for practical use due to its high performance overhead and cost. The Adaptive Reconfigurable Mobile Objects of Reliability (Armor) middleware architecture offers a scalable low-overhead way to provide high-dependability services to applications. It uses coordinated multithreaded processes to manage redundant resources across interconnected nodes, detect errors in user applications and infrastructural components, and provide failure recovery. The authors describe their experiences and lessons learned in deploying Armor in several diverse fields.
What problem does this paper attempt to address?