Rebecca Isaacs, Amazon
Distributed systems have complicated and unpredictable behaviors, driven by the dynamics of the arriving workload, faults, failures, and interactions with other systems. One of the big challenges for operators is to ensure that their system is resilient, meaning that it can tolerate inevitable stressors like overload or hardware failure, and either continue operating, perhaps in a degraded state, or fail gracefully. We have found that simple models and analytical techniques help to ensure the desired level of resilience by demystifying the behavioral dynamics and the trade-offs of complex distributed systems. In this talk I will describe work we’ve done at AWS over the last couple of years to understand the vulnerability of services to metastable failure (congestive collapse), and to reason about the tradeoff between server protection and client availability in retry policies.

Rebecca is a senior principal scientist at AWS, where she is part of the DynamoDB organisation, working on resilience and performance. Prior to AWS, she worked at Twitter and Google, largely focused on all aspects of distributed tracing, from trace production through to novel uses of aggregate trace analysis. This followed over a decade at Microsoft Research doing research broadly in the area of performance analysis of distributed and concurrent systems.

author = {Rebecca Isaacs},
title = {Analysis for Better Resilience},
year = {2026},
address = {Seattle, WA},
publisher = {USENIX Association},
month = jul
}