Fail-slow at scale: Evidence of hardware performance faults in large production systems

56Citations
Citations of this article
32Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Fail-slow hardware is an under-studied failure mode. We present a study of 114 reports of fail-slow hardware incidents, collected from large-scale cluster deployments in 14 institutions. We show that all hardware types such as disk, SSD, CPU, memory, and network components can exhibit performance faults. We made several important observations such as faults convert from one form to another, the cascading root causes and impacts can be long, and fail-slow faults can have varying symptoms. From this study, we make suggestions to vendors, operators, and systems designers.

Cite

CITATION STYLE

APA

Gunawi, H. S., Suminto, R. O., Sears, R., Golliher, C., Sundararaman, S., Lin, X., … Li, H. (2018). Fail-slow at scale: Evidence of hardware performance faults in large production systems. ACM Transactions on Storage, 14(3). https://doi.org/10.1145/3242086

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free