FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production Monitoring

4Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

This paper presents Meta's FBDetect system, which advances the state of the art in performance regression detection by catching regressions as small as 0.005% in noisy production environments. FBDetect monitors around 800,000 time series covering various types of metrics (e.g., throughput, latency, CPU and memory usage) to detect regressions caused by code or configuration changes in hundreds of services running on millions of servers. FBDetect introduces advanced techniques to capture stack traces fleet-wide, measure fine-grained subroutine-level performance differences, filter out deceptive false-positive regressions, deduplicate correlated regressions, and analyze root causes. Beyond these individual techniques, a key strength of FBDetect over prior work is its battle-tested robustness, proven by seven years of production use, and each year catching regressions that would have wasted millions of servers if left undetected.

Cite

CITATION STYLE

APA

Yoon, D. Y., Wang, Y., Yu, M., Huang, E., Jones, J. I., Kukkadapu, A., … Tang, C. (2024). FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production Monitoring. In SOSP 2024 - Proceedings of the 2024 ACM SIGOPS 30th Symposium on Operating Systems Principles (pp. 522–540). Association for Computing Machinery, Inc. https://doi.org/10.1145/3694715.3695977

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free