🎙️ Episode 2104:11October 26, 2025

AWS Outages

Listen to this episode

AI-generated discussion by Alex and Jamie

About this episode

Alex and Jamie unpack AWS Outages — what shipped, why it matters, and how engineers can put it to work today. New episodes weekly.

Transcript

Welcome back, folks, to your favorite deep dive into all things tech, the nerd-level Tech AI cast. I'm Alex, bringing the wisdom of the ancients, or at least what feels ancient in tech years. And I'm Jamie, here to ask the questions you're all thinking and maybe crack a joke or two along the way. Today, we're talking about a topic that's as thrilling as finding out your favorite show got renewed for another season, AWS outages. Oh, the drama, the suspense. It's like the season finale cliffhanger nobody asked for. AWS, or Amazon Web Services, is the giant whose shoulders the internet stands on. But even giants can stumble. And when they do, it seems like the whole internet feels the earthquake. So what happened with these AWS outages? Let's start with June 2023. AWS's US East 1 region, think of it as the New York City of cloud regions, had a bad day, a really bad day. For almost four hours, services like the Boston Globe, Southwest Airlines, and even Taco Bell's app were thrown into chaos. Wait, you're telling me my taco order got lost in the cloud? That's a new one. Exactly. But it wasn't just any cloud problem. It was a hidden software defect in AWS Lambda, which is like the kitchen where all the internet's orders are cooked up. This defect had been lying in wait, a ticking time bomb, until conditions were just right, or just wrong, in this case. So it was a sneaky bug. How did they fix it? AWS engineers were on it pretty quickly. They put a stop to new Lambda invocations, hitting the buggy code path, and rolled out a permanent fix. It was like rebooting the kitchen without stopping all the cooks. Fast forward to October 2025, and it's deja vu all over again. This time, it was a DNS race condition in DynamoDB. Could you break that down for me? Imagine two automated systems both trying to change the same light bulb at the same time. They get in each other's way, and the light goes out. That's essentially what happened. But with DNS, the internet's address book. When services couldn't find DynamoDB, they essentially couldn't find their data. So no data, no service. I'm seeing a pattern here. Right. The cascade effect is real. And the recovery was complicated by a thundering herd problem. Imagine trying to restart a stadium's worth of services all at once. They had to throttle the restarts to get everything back online, without crashing again. This sounds like a lot of tech headaches. How is AWS making sure this doesn't keep happening? They've expanded with new regions, making the internet less reliant on any single one. They've also improved their monitoring with generative AI observability for CloudWatch, which is a bit like having a super smart watchdog. But even with all these improvements, we've seen that AWS, like any complex system, can fail. What's the takeaway for businesses relying on cloud services? Diversify and prepare. Use multiple regions, maybe even multiple cloud providers. And always have a plan for when things go south, because at some point they might. It's been quite the journey today, from lost tacos to global outages. Any final thoughts, Alex? Just that in the world of tech, change is the only constant. And resilience is a dance, not just a set of backups. Wise words to live by. Well that's all the time we have for today's episode of Nerd Level Tech AI Cast. Thanks for tuning in, and remember, when in doubt, reboot. But maybe also have a backup plan. And don't forget to follow us on your favorite podcast platform for more tech deep dives. Until next time, keep your data safe and your tacos closer. Bye everyone. Catch you on the next wave of the tech tsunami.