NASA Logo

NTRS

NTRS - NASA Technical Reports Server

Press Enter or click the Search button to begin your search.

Back to Results
Checkpoint-based forward recovery using lookahead execution and rollback validation in parallel and distributed systemsThis thesis studies a forward recovery strategy using checkpointing and optimistic execution in parallel and distributed systems. The approach uses replicated tasks executing on different processors for forwared recovery and checkpoint comparison for error detection. To reduce overall redundancy, this approach employs a lower static redundancy in the common error-free situation to detect error than the standard N Module Redundancy scheme (NMR) does to mask off errors. For the rare occurrence of an error, this approach uses some extra redundancy for recovery. To reduce the run-time recovery overhead, look-ahead processes are used to advance computation speculatively and a rollback process is used to produce a diagnosis for correct look-ahead processes without rollback of the whole system. Both analytical and experimental evaluation have shown that this strategy can provide a nearly error-free execution time even under faults with a lower average redundancy than NMR.
Document ID
19940025365
Acquisition Source
Legacy CDMS
Document Type
Thesis/Dissertation
Authors
Long, Junsheng
(Illinois Univ. Urbana-Champaign, IL, United States)
Date Acquired
September 6, 2013
Publication Date
January 28, 1994
Subject Category
Computer Systems
Report/Patent Number
UILU-ENG-94-2201
CRHC-94-01
NAS 1.26:195760
NASA-CR-195760
Report Number: UILU-ENG-94-2201
Report Number: CRHC-94-01
Report Number: NAS 1.26:195760
Report Number: NASA-CR-195760
Accession Number
94N29869
Funding Number(s)
CONTRACT_GRANT: NAG1-613
CONTRACT_GRANT: N00014-91-J-1283
Distribution Limits
Public
Copyright
Work of the US Gov. Public Use Permitted.
No Preview Available