Vlad LeyberovGuest
Alexa GriffithHost
So every single one of these notifications, like Sarah liked your photo or Mike commented on your post, had to reach the right device at the right time without delay.

Now imagine one configuration change, one retry bug, one cascading failure, Suddenly, billions of daily notifications don't arrive.

Users wonder, what's going on? Is this app broken? So someone has to keep that running.

Someone has to make the decision like, do we intentionally take the system down 100% to fix it faster? Or


These systems are the ones we use every single day, systems where even a small mistake can have massive global consequences.

It's quite difficult to comprehend, especially at large scale, how cascading failures could propagate throughout the system.

What I've observed over the years of working in large distributed systems, that their behavior is unpredictable.

Even a little change, if we don't fully understand how cascading failures propagate, will have absolutely catastrophic consequences.

This episode is a conversation about what it actually takes to keep the internet running when failure just simply is not an option.
Read the full transcript.
Create an account to read the whole episode, search across every transcript, and follow the shows you care about.

So every single one of these notifications, like Sarah liked your photo or Mike commented on your post, had to reach the right device at the right time without delay.

Now imagine one configuration change, one retry bug, one cascading failure, Suddenly, billions of daily notifications don't arrive.

Users wonder, what's going on? Is this app broken? So someone has to keep that running.

Someone has to make the decision like, do we intentionally take the system down 100% to fix it faster? Or


These systems are the ones we use every single day, systems where even a small mistake can have massive global consequences.

It's quite difficult to comprehend, especially at large scale, how cascading failures could propagate throughout the system.

What I've observed over the years of working in large distributed systems, that their behavior is unpredictable.

Even a little change, if we don't fully understand how cascading failures propagate, will have absolutely catastrophic consequences.

This episode is a conversation about what it actually takes to keep the internet running when failure just simply is not an option.