Benjamin Leonhardi, Meta Platforms; Evangelia Kalyvianaki, University of Cambridge; Yang Wang, Meta Platforms and The Ohio State University; Abdelrahman Adam, Agshin Nabiyev, Aleks Shirokov, Amitav Mohanty, Daniil Balenko, Elaine Zhao, and Essam Ewaisha, Meta Platforms; Hongbo Dong, NexGeMM LLC; Igor Marnat, Lev Novikov, and Min Zeng, Meta Platforms; Steven Shingler, Independent Researcher; Timofey Durakov, Wiliam de Abreu Pinho, Ben Christensen, Mayank Pundir, and Kaushik Veeraraghavan, Meta Platforms
Maintenance is a fundamental operation in datacenters to ensure that hardware and software operate correctly, efficiently and use up-to-date versions. We present Meta’s maintenance system in production over the last five years that provides continuous support to tens of thousands of services running on our fleet of millions of servers and seamlessly orchestrates this process. To our knowledge, this is the first paper to discuss predictable maintenance at scale.
The key challenge is how to minimize the capacity buffer—servers reserved to absorb capacity loss caused by maintenance—while providing a predictable latency to maintenance operations. This paper presents a series of strategies and techniques we use to accomplish this goal, such as aligning maintenance with fault domains, placing hardware evenly across fault domains, a maintenance contract among participating parties, etc. Indicatively, we observe that these techniques have helped us reduce the size of the capacity buffer by about 15% in one quarter of 2025 and allowed us to perform a fleet-wide deployment under targeted SLOs (e.g., 45 days for a new OS, 90 days for a new firmware).
OSDI '26 Open Access Sponsored by
King Abdullah University of Science and Technology (KAUST)
Open Access Media
USENIX is committed to Open Access to the research presented at our events. Papers and proceedings are freely available to everyone once the event begins. Any video, audio, and/or slides that are posted after the event are also free and open to everyone. Support USENIX and our commitment to Open Access.

author = {Benjamin Leonhardi and Evangelia Kalyvianaki and Yang Wang and Abdelrahman Adam and Agshin Nabiyev and Aleks Shirokov and Amitav Mohanty and Daniil Balenko and Elaine Zhao and Essam Ewaisha and Hongbo Dong and Igor Marnat and Lev Novikov and Min Zeng and Steven Shingler and Timofey Durakov and Wiliam de Abreu Pinho and Ben Christensen and Mayank Pundir and Kaushik Veeraraghavan},
title = {{PIMS}: {Fleet-Wide} Datacenter Maintenance with Minimal Capacity Buffer and Predictable Latency (Operational Systems)},
booktitle = {20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26)},
year = {2026},
isbn = {978-1-939133-55-7},
address = {Seattle, WA},
pages = {2169--2185},
url = {https://www.usenix.org/conference/osdi26/presentation/leonhardi},
publisher = {USENIX Association},
month = jul
}