Transcript

Sam Newman: I'm talking to you today about this concept called progressive collapse, which I came across while doing research for my latest book. I came across this while looking at the wider concept and wider space of resilience engineering, and lessons and things that we can learn from topics and from domains that are not around computing. I learned about the story of Ronan Point. This is a tower block that shortly after it was opened in 1968, suffered a partial collapse. It was not a good time. This is in Canning Town. You can get a sense of the scale of what happened in this wide shot. Only four people died. I say only four people died. It was a miracle it wasn't more. The only reason more people didn't die was because of the hour at which this particular collapse happened. It was quarter to 6:00 in the morning.

There's basically a corner gone. You can see the mirroring sister block on the right-hand side here. Most of those rooms that collapsed were living rooms that early in the morning, very few people were up. We're going to be talking about what happened to Ronan Point and this concept called progressive collapse, which comes from the space of civil engineering. Also going to take a look at how we can take that metaphor and use that to reason about our own digital systems, when things go wrong in the distributed digital systems that we're all creating and hopefully trying to make more resilient. Taking those lessons that we can learn from the building industry, what are the techniques that we can use to stop progressive collapses happening in our own system? I'm going to take you into a bit more detail about Ronan Point, but also a couple of other examples of system failure, one of which I was intimately involved with and one of which may have impacted many of you last year.

Examples of Progressive Collapses

What happened to Ronan Point? These buildings were built quick. This is post-war, although this is 1960s, there was still an awful lot of London, especially in the Canning Town area, that was still rubble after the Blitz. There was a boom in population, we needed more housing, and so a lot of these blocks went up quick. We'll come back to the building methods used a bit later on. It was early in the morning, resident Mrs. Ivy Hodge got up. First thing any self-respecting British person does, especially in the late '60s, was make tea. She put a stove on, went to heat up some water, and there was a gas explosion. The gas explosion was caused by just a faulty nut around where the oven was connected to the main gas supply. This was a survivable gas explosion. I mean that quite literally in terms of it wasn't a great time for Mrs.

Hodge. She was blown across the room, knocked unconscious. This is what's left of her kitchen. She came to in a puddle of water from the saucepan that was blown off the top. This was not a big gas explosion in the grand scheme of things. Still, not an ideal scenario for first thing in the morning. You can actually see some of the problems start to occur beyond that Mrs. Ivy Hodge was in there somewhere. When you look, you can see through the side of the building now. This is where our problem gets a lot worse. This gives you a plan idea of what happened. You can see where the sink was. You can see the gas cooker was knocked over. Then you see where the building used to be on the left-hand side, that her bedroom was sheared off. What had happened was that the explosion blew out the outer wall, which happened to be load-bearing.

The four floors above collapsed down, and that created a concertina effect, taking down that corner of the building. You can see where the explosion happened. We might assume that the impact would have been even greater if the explosion happened further down the building. This is what's known as a progressive collapse. A progressive collapse describes a situation where a small failure results in a significant collapse in the wider system. The initial triggering thing in isolation looks small, somewhat innocent, but it ends up cascading to cause a much bigger impact. We're all familiar with this idea. I'm sure some of you have either done this experiment or seen versions of this where you can start off with a little small domino hitting a bigger domino. I've seen examples where they're knocking over dominoes that are like over a ton in weight. The differences with this kind of system is that we can see all the moving parts and we can reason about it. The progressive collapses that happen in our digital distributed systems are not as easy to necessarily understand.

We're going to take a look at two other examples of progressive collapses that occurred. Firstly, we'll take a look at the us-east-1 outage in October of last year, and secondly, a website selling used cars. This is my fault, but we'll come back to that. Let's talk about what happened with AWS. Actually, out of interest, were you or any of your systems impacted by this outage? A surprising number of UK banks were impacted by this, which doesn't really make a lot of sense considering there is an AWS region here, but we'll maybe come back to that later. Fundamentally, what happened was that a subsystem of DynamoDB was tasked with updating DNS routes within the AWS infrastructure. This was to allow for things like provisioning, load distribution and the like. This had an error, which we'll get to the details of later, and this caused DNS routes in Route 53 to be deleted.

This in turn had a cascading effect because other services then that relied on this DNS route started to fail. Network load balancers started going down, compute started failing, queue stopped working, and EKS also failed. This is all compounded because a lot of these lower order services in AWS are used by other services in AWS. When those network load balancers start going, all hell breaks loose. A lot of the us-east-1 services started failing. This in turn took customer-facing products and companies offline. Amazon's own Alexa stopped working. Ring, Slack, Snapchat, Zoom, Shopify all suffered either partial or total outages as a result of these particular failures. People's alarms didn't go off, so they were late to work. It seems like a good excuse to have, though, if you are late for work. "I'm sorry, a cloud region went down." People's smart locks would not unlock or would not lock.

At least one person on Reddit claiming they were locked in their house. More humorously, people who have all the money in the world can buy a $2,000 smart mattress, found their mattresses were stuck in certain positions, were stuck heating or cooling, as the case may be. The one I liked the most was if you have $600 bucks, you can buy a smart litter tray. These smart litter trays also stopped working. Don't worry, the cats were not trapped. We had these interesting rippling effects that occurred.

Let's actually go a bit deeper into what happened. We're getting this information from AWS's own incident report. This component of DynamoDB that updates the DNS routes, there's two main parts of it. Think of a DNS planner, which worked out which DNS entries need to be updated. It would create plans, which makes sense. Then you had a number of DNS enactors, and their job was to enact the plan. Naming stuff, it turns out, is quite easy. What happened was a DNS enactor would pick up a plan and apply those changes. There were multiple copies of the DNS enactor for resiliency, which really ended up causing this problem. We'll come back to that later on. The issue was on the particular day of the incident, was one of these plans took a long time to implement, much longer than normal, was caused by a host of underlying issues.

This plan had not completed, a new plan got created, a different DNS enactor picked up that other plan, it completed its work very quickly. Unfortunately, then the old plan completed and wiped out all the routes. There is more detail here in the incident report. I'm glossing over a few details. This was a classic race condition resulting in all those DNS entries being wiped out.

What happened with the used car site? What about this kind of failure? This was a website, it still exists, I was involved in many years ago, was for helping you find used cars, motorbikes, caravans, and the works. It has historically been a bunch of websites based around those particular verticals. We were in the process of combining all of those verticals into a single application. This is actually a classic Strangler Fig application. In fact, this is the case study from Martin's original write-up. Our system, which is codenamed Sauron, was running on 10 servers. Typically, even at peak load, we would expect to have something like between 30 and 60 concurrent requests at any given point in time. Some of that traffic was being served directly by Sauron. Other calls coming in, though, were being passed on to the downstream existing legacy applications. On the day the incident occurred, these nodes started running a little bit hot.

They went from handling between 30 to 60 concurrent requests to handling over 800. This is a Java stack, native threading, looks interesting, we're going back to green threads again. That meant each of those concurrent requests became an operating system thread. The entire CPU was saturated as it was spending all of its time trying to sequence between those threads. This took the whole system down. It went down quick. What was the issue? We had all these requests coming in. It wasn't a denial-of-service attack. We eliminated that pretty quickly. Turned out the issue was that one of the downstream sites started behaving in a suboptimal way. It would let you establish a connection, and it would just hang. It would just hang forever. We were timing out way too leniently. We were taking 30 seconds before we timed out those calls. This basically resulted in the connection pool that we were using to make calls becoming completely exhausted.

Now we had a thread pool issue. Because that connection pool was exhausted, any other calls coming in that were for motorbikes or for jet skis were unable to get a workup. Those calls themselves were also blocking. This was made worse because when the website hung, what would people do? They hit refresh. They would just hit refresh. Then the pages wouldn't load. They hit refresh again. The pages wouldn't load. They kept hitting refresh. We weren't timing out on the calls coming in. Not ideal. There were a bunch of different things that we did to resolve this issue going forward. I'm going to take you through some of them and the ones that we missed at the time.

How to Mitigate a Progressive Collapse

These are three examples we looked at. We've got Ronan Point. We've got AWS. We've got the used car site. How do we mitigate for a progressive collapse? We can't eliminate the possibility. What things can we do to mitigate against these things happening in our own systems? I want to preface this by saying that focusing on a single root cause is highly problematic. Note to whichever marketing team works for one of the vendors that keeps talking about finding the root cause. No. Bad vendor. Thinking about a single root cause is deeply problematic. At the very least, you should use the word root causes. In general, I try and avoid this because it makes us think overly simplistically about the world. You might be in a region which is troubled by forest fires. I lived for many years in Australia. You might think there was a big fire.

Not good. It was caused because somebody dropped a match. That's the root cause. Let's stop people dropping matches. That might be a good idea, but it's not by itself sufficient. In Australia, for example, you're not allowed open fires in bush areas. You're not allowed to use things like fireworks either because they cause fires. Doing that is useful, but it's not by itself sufficient. They do things like they clear the brush and the undergrowth. They do controlled burning to reduce the load within the forest as well, because they recognize that although they can try and remove the thing that maybe causes the fire initially, you can't remove all causes of fire. Lightning is a thing. Other things are needed to be done to reduce the impact of a forest fire if it occurs. The problem with this is that a lot of this viewpoint about finding the root cause still permeates our industry.

I often think that IT is still stuck in the world of traditional safety management. This is a fixation on the idea that we just reduce adverse events and that makes things better. We try and stop anything bad from ever happening. The problem is as our systems get bigger and our systems become more complicated, we cannot stop all the bad things happening. In any case, a lot of the bad things are entirely out of our control. This is why in the world of resilience, we've gone from this world of traditional safety management and have moved over to the world of resilience engineering, which is a big topic and it goes far further. We can think of resilience engineering as accepting that things can go wrong and still trying to mitigate them. However, accepting that and trying to make sure as many things go right as possible. This is a very different mindset. Accepting that things might go wrong, but knowing how to mitigate those things when they occur.

In that spirit, what can we do to mitigate a progressive collapse? Of course, at this point, I go to your friend and mine, NIST. I popped over to NIST, which is replete with very interesting PDFs on all kinds of topics. We have this report on practices reducing the potential for progressive collapse in buildings, not in distributed systems. That's the only flaw with this particular report. It is 216 pages. I did not read all of them. I put it into NotebookLM and I asked NotebookLM questions. You can download it. There's actually lots of really interesting advice in there. I was basically able to synthesize the three key ways that you can mitigate for a progressive collapse. I'll talk to you about how those were used in the context of buildings. I'll also talk to you about how those three types of mitigations can be applied in our digital systems.

We're going to look at the importance of reducing hazards, of strengthening components in our system, and reducing the interconnection in our systems. We're going to look at each of these three things in turn. If you're trying to try and mitigate progressive collapse in your own environment, you probably want to do something in all of these areas, if you can.

1. Reduce Hazards

Let's talk about hazard reduction. In the wake of Ronan Point, what could we do to reduce hazards? Immediately in the aftermath of the incident, it was suggested to stop allowing gas to be installed in high-rise buildings. A lot of these buildings were going up. There's over 1,000 buildings that have the same building system as this across the country. Just don't install gas. That would lose at least that inciting element. In those situations where gas was already installed, it was suggested that we should improve ventilation. This avoids gas building up, but also makes it more likely that other people are going to smell the gas and alert the authorities. Any of you who have spent any time living in the UK know ventilation and where we live tend not to line up very well. We all have damp issues around the winter, don't we? Nonetheless, it seems like a sensible thing to do to reduce the hazard of a gas explosion.

What about AWS? We can look into at least what AWS talked about doing again from their incident report. Very sensibly, they said, we've clearly got an issue with our automated DNS management, so while we work it out, we're turning it off. We've got a fallback. We're just shutting the thing down. Very sensible. Glad they did that. In addition, they were going to put some improved testing around this particular component to ensure that if a subsequent issue came up again in the future, they'd catch this before it went out to production.

We move over then to the used car site, and context here when we're considering mitigations is always important. We thought, what can we do to strengthen this caravan component to stop it failing in the way that it's failing? We could go in, we could try and fix the code and solve this problem. This was a legacy application that we were going to retire. It also represented only a fraction of the revenue anyway. This was not a critical business import to us. There was no real interest in actually fixing it. There was no real interest in strengthening the component itself because it didn't make sense from a business point of view. What we did at least do is put some improved monitoring around it to try and pick up these issues if they occurred in the future. What I found out about years later, but I wish I'd known at the time, was we could have put some load shedding into this component to help reduce the hazard.

The hazard here of Sauron as a whole was that request count, the resource saturation. The resource saturation was caused by having too many requests. We could have just shedded the load. With load shedding, you basically say, I can only handle 30, 60, 50, whatever how many requests you can handle. Anything over and above that, I'm just going to drop on the floor. That would have kept our system up and running. If we'd applied that load shedding protection to the Sauron application, Sauron would have carried on working. It's always nice, wish I had a time machine right about now.

2. Strengthen Components

We've reduced our hazards, so how can we strengthen our components? How can we reduce the chances of those components failing in their entirety? Strengthening things is a good idea. We strengthen concrete with steel in it. That makes it stronger. This was one of the major problems with Ronan Point. When it was built, there were not enough people with the skills required to build buildings how we normally build buildings. A system was selected that would allow lower skilled people to build high-rise buildings quickly. We came up with a system that allows lower skilled people to build high-rise buildings quickly. What could go wrong? In the immediate aftermath of this, it's like, maybe we should look at whether or not we should be doing this, full stop. Because it turned out that there were some significant issues in how this building had been put together. These pre-fabricated sections were brought onto site and they were slotted into each other like Lego.

You had these steel ties that would overlap, and then the joins between the panels were supposed to be filled with concrete. Yes, concrete has very particular material properties. The lower skilled people that hadn't had very good training and where there was no real oversight had made some different decisions on the fly and thought, we could use concrete or we could use newspaper and cigarette packets. I don't know if you're aware, they have different tensile properties to concrete. This was only found because there was an almost forensic deconstruction of the building later on that discovered this issue. This, as a result, strengthened and improved building regulations in the UK, had a lot more oversight. They also started instituting things like tool chest talks, almost like stand-ups at the beginning of the working day for construction, recognizing that whether or not the large panel system itself was flawed, it was not implemented the way it should have been. This actually changed the building industry in the UK and beyond, because we all got to learn from this.

How do we strengthen a service? One way we can strengthen a service is by improving the redundancy in the components that we use to operate that service. We're aware, for example, that failure of a service instance is a thing, and so we decide to have multiple copies of a service instance and stick it behind a load balancer, because the idea is that if one of those instances fails, our service can continue to operate. The service as a whole has been strengthened through the provisioning of redundant resources. This is so much an obvious known issue. Let's come back to caravans, for example. Could we have strengthened our components there by running redundant copies? Yes, absolutely. The caravan site was actually a read-only application. Running multiple nodes would absolutely have made sense. Do we want to spend money running redundant nodes of a service that we're planning to retire and that represents a fraction of our traffic and our revenue?

No. Financially, this was absolutely a non-starter at the time. There were some other things we're going to have to do here. What about in AWS, strengthening the component? Actually, there was some redundancy already within the existing DNS solution. Those multiple DNS enactors was a source of redundancy. Weirdly, of course, that actually also caused the race condition. Strengthening the component in the context of AWS specifically would maybe involve fixing their race condition. Then we have to look broader. All those people that are impacted by their use of AWS. If you think about the cloud as the component, how do I strengthen my cloud? If I am the CEO of Eight Sleep, or whoever the smart litter box is, what else could I do? In the aftermath of the region failure, everyone started saying, you need multiple sites. We need to be running across multiple clouds. The companies themselves were a bit quiet about whether or not they were going to do this, but everybody on Hacker News thought this was a good idea. Clearly, it must be a good idea.

At first glance, this does seem to make sense. If you look at running on AWS, we have the concept of multiple cloud regions. This gives you not only being able to operate in different countries, but just separating out your blast radius. We could. Why not run our system in multiple sites, across multiple regions? Even when you decide you're going to do that, you've got to decide, what technique are we going to use? There's a whole plethora of different ways that you can run your system in a multi-site setup. This is traditional DR type stuff. We can start with a classic backup and restore mechanism. We could back up our data from us-east-1, and if something goes wrong, we can reconstitute that somewhere else. Guess what? That takes ages. Also, if it's going to work, you have to have invested a lot in automating your infrastructure.

Plus, you also need to know you're going to get the infrastructure somewhere else. Everyone should be doing backups anyway. This is typically a good starting point if you're looking at a multi-site solution. We've got things like the pilot light setup. What the pilot light setup is trying to do is reduce your window of data loss. If you back up data once every hour and you reconstitute your system from those backups, you can lose up to an hour worth of data. With a pilot light setup, you're more constantly replicating data from your primary site to your failover site, but you don't actually have any infrastructure running to serve your traffic. The idea being that you shrink that window of data loss. The idea then is when you have a failover mode, you can spin up the infrastructure. It's about keeping your costs low in a dynamic provisioned environment.

You've then got variations, the warm failover. You have some infrastructure provisions so that you can fail over some traffic or some critical paths. Again, it's all about balancing cost. Of course, the holy grail of this is let's just go multi-site then. Let's have two sites running and allow traffic to be sent to either site because that's going to eliminate my data loss window. It never eliminates it entirely. Also, I'm not going to have any downtime assuming one of the sites has enough capacity to handle all of the load.

Of course, this looks great. The problem here is the complexity. I need to be really clear about this in case you've had conversations with any vendors. If anybody tells you that two-way data synchronization is really easy or simple, they are trying to sell you something. Unless you started off building your application with this in mind, retrofitting this into an existing system is not trivial. Do your research before you think about whether or not this is for you. Fundamentally, when we start looking at these multi-region, multi-site options, as we come up with these patterns, we are reducing our downtime, we're reducing our data loss windows, but we are increasing cost and complexity. As we go from left to right, we're getting better from a point of view of downtime and data loss, but we are getting worse from a point of view of cost and complexity.

A lot of the time, this comes down to money. Why weren't more of those big-name companies running a multi-site setup? Money, in part. Those smart people at these companies, do you think they didn't know that they were based in one region? Of course, they knew. They absolutely knew, and they had made a decision not to have a hot failover. There are reasons behind that. Some of that might be down to cost. Are they going to charge you more for the work required to do this, when these failures occur very rarely? Maybe not. There are other elements, though, to doing multi-site that we'll come back to a little bit later on. I got a lot of good guidance around some of these patterns from this really nice write-up of the Well-Architected Framework, AWS. These patterns are fairly generic. They don't obviously talk about multi-cloud. Weirdly, that AWS is not talking about multi-cloud, but there is some good stuff in here. You might also be thinking about data sovereignty. If your concern is around wanting to run multi-region or multi-provider for data sovereignty reasons, that is completely legitimate. That's not what we're talking about today. We will come back to that concept, though, a little bit later on.

3. Reduce Interconnection

Let's look at our third way of mitigating our problems, which is reducing interconnection. If we think about that nice little conveyor belt of dominoes, if you want to stop all the dominoes getting knocked down, you can just go and remove one of the dominoes. You get a bit of a collapse, but it stops. Reducing the interconnection in our system can be a great way to make sure that a problem in one area doesn't cascade through the entire system. Within Ronan Point, the problem was that there weren't alternative load-bearing parts. The load-bearing was done on the outside wall. Once that wall was removed, the rest of the floors came crashing down. By establishing alternative paths through which the load could travel, you avoid that problem. A load-bearing wall being removed isn't going to result in a catastrophic incident. In our IT systems, in our distributed systems, we'll talk a lot about bulkheads.

Bulkheads is a metaphor that comes from shipping. This is a picture of a submarine. If you hit a rock, water starts pouring in. You can close that compartment, assuming you're not in the compartment. If you are in the compartment, bad news. The idea is that the ship remains seaworthy. Famously, with at least one theory that the reason the Titanic sunk was that its bulkhead design was actually quite seriously flawed. The idea here is that we compartmentalize the failure. We reduce the interconnection. If you're the owner of a smart sandbox, how do we decrease our interconnection? How do we reduce the interconnection? How about having either a local fallback, or even better, being local first? The CEO of the smart mattress company, I think they're called Eight Sleep, said, we are going to have a local fallback. Local fallback's fine. This is effectively a bulkhead that closes when there's an issue, which is what bulkheads do. Local first might be even better than that, but that's a start.

What about multi-cloud? Because with multi-cloud, my concern is what if a cloud vendor fails? The challenge always, when we start looking at multi-cloud around reducing, because that's me reducing my connection to a particular vendor. The problem with a multi-cloud approach is you're taking all that complexity that we've already looked at in terms of multi-site, but now we're overlying the need for you to have skills and expertise in multiple cloud vendors, or for you to curate effectively a commoditized layer. I know Martin talked about this idea that Kubernetes is effectively a commodity layer for compute, and he's right. We also know that's not the whole story. You can't just pick up an application here and stick it on any old Kubernetes cluster. There's a lot more stuff that has to go in. That's investment you have to put into, beyond the fact that operating EKS is quite different to operating AKS.

All of this gets worse if you throw multi-cloud into it. If you're interested in looking at multi-cloud, it's really important that you see this through the lens of hedging your risk against a supplier failing, not a site failing. This might be the risk you're in. You're reducing your interconnection on a particular vendor. Again, there could be great reasons to do this around things like technical sovereignty or your concerns around the PATRIOT Act or the CLOUD Act, for example.

There's this annoying problem though, that even with multi-site setups or multi-cloud setups, that actually we can cause problems around interconnection. We are replicating data across those sites. That can actually cause problems. There's a reason why there are very few cross-regional cloud-based services provided by any of the big cloud vendors, because they recognize that any connection between sites becomes a potential infection vector. Let me simplify that down in a way. If you deploy the same code in vendor A as you do in vendor B, your code itself is a form of interconnection. A bug that takes down your system on AWS could just as likely take down your system on Azure or your own on-prem system. This is why some firms go even further to reduce their interconnection. We can look at someone like Monzo. They have a banking license. They actually have to be certified.

They have to have the regulator come in and approve what they're doing. One of the things they have to provide is continuity of services. Rather than taking the services of a British bank and running it in a U.S. cloud region and thinking, that's me done then, they've gone a bit further. They've said, yes, we're a bit worried about having an issue with a vendor or a site going down, but we're also just as worried about a bug in our code causing a problem in other people's systems. They've created a stand-in system that is a cleanroom re-implementation of the critical banking functionality. No shared code, very different architecture, stripped down, very simple, that runs in GCP, the core system runs in AWS. They can constantly validate their stand-in so when you go and use Monzo, you might be opted in to using the new site.

What about with our used car site? We started doing things like separating our connection pools. We had multiple connection pools for each downstream server. We also started using circuit breakers. Many of you know about circuit breakers. The idea is that after a certain number of calls, we blow that circuit breaker open and make sure the traffic doesn't happen. You can think of a circuit breaker as a way of triggering a bulkhead on demand. When that circuit breaker is open, we've reached a threshold, we can now move back to failing fast. There are a bunch of issues around circuit breakers. They aren't simple things. They can cause more problems in your system. I'd thoroughly recommend if you're using circuit breakers to read these two posts by Marc Brooker. He points out some of the challenges around them and actually recommends some simple mitigations like only using circuit breakers around retries or else considering using things like token buckets, which can be much more elegant ways of handling this.

Actually, adding circuit breakers can make other problems worse. This is an inherent paradox we have in our systems. As we increase the system's complexity to handle issues, we can end up introducing more sources of failure. This is a real problem. We think we're worried about a service misbehaving. We have a circuit breaker that opens when that service misbehaves. Fantastic, that's great. Then we start seeing problems like it makes partial failure worse. If I've got a service endpoint and some of the functionality works really well and some of the functionality is erroring, the functionality that's erroring could cause a circuit breaker to open. Now the functionality that was working can't be used. We also have that general boom and bust problem. All of these clients start tripping their circuit breakers at once. Our service goes from being overloaded to not loaded at all. We can get those floods coming through. Circuit breakers being incorrectly configured can cause retry storms, which can in turn bring systems down.

We think, I want to make sure my service doesn't fail. The way I'm going to handle this, Sam, is by having multiple copies of my service. I'm going to do that with Kubernetes. What does that look like? This is your system on drugs. This is the stacks that we now run, and every single layer of these stacks makes sense in and of themselves. As we strive to make our systems more resilient, we can actually end up increasing the complexity of our applications to the point where we introduce new sources of problems, new sources of errors. This is something that David Woods has talked about a lot. This is, to an extent, not solvable. We just might have to be aware of it. Sometimes the right decision is to say, that's not for us. Yes, we can do it, but we're going to make a conscious decision not to.

It is completely ok for you to make a conscious decision that multi-cloud is not right for you because you're worried about the ongoing complexity you're going to have to manage, or that maybe the increased interconnection of those systems might cause more issues down the line. Think about the AWS outage. They decided to run multiple copies of that DNS enactor. Why? Because they knew that availability zones could fail. They ran multiple copies. If they only ran one DNS enactor, this wouldn't have failed in this particular scenario. There wouldn't have been a race condition. It wouldn't have been possible for that race condition to exist because plans would have ended up being executed sequentially. That complexity they added allowed the race condition to occur, which ended up bringing that system down. It's really annoying, isn't it, when you get into these sorts of things? My urging for all of you is not to say that you shouldn't do this stuff, but you make a conscious decision.

There's a lot of learning by rote that goes on. We're often time poor. We've got a lot of stuff going on. We just do things because other people have done them or we saw them in a conference talk, but we don't necessarily understand the realities of what that causes. Finding time every now and then to have a chat as a team and say, is this the right thing for us to do? Is a good idea. No, sticking circuit breakers everywhere isn't necessarily going to solve your problems, might just give you some new ones.

Summary

I haven't tried to boil the ocean with this talk. There's a huge amount more we can talk about in terms of concrete practices to mitigate progressive collapse. The reason I did this talk is I always find it interesting to look at the intersections between things, to look at different industries, different domains, and see what we can learn from them. Metaphors like the bulkhead and circuit breakers, these are concepts that come from industries not like ours, but they are now key parts of how we think about our own system resilience. If we're thinking about progressive collapse, progressive collapse is a concept that comes from the building industry, but it's still things that we can learn from it. These progressive collapses seem innocent when they first start, but the results can be significant. You've probably all experienced something like this, that small failure which results in a collapse in the wider system.

It's not a fun time. What can we do if we want to mitigate a progressive collapse? Number one, reduce hazards. Stopping things breaking is still a good use of your time. It's not by itself sufficient. Two, strengthen the components, make them more resilient when those hazards still occur. Number three, we can reduce the interconnection between things. We've looked at things like reducing hazards. That load shedding on Sauron would have reduced the hazard. It would have reduced the load on that component and stopped it from being brought down. We've looked at ways to strengthen components, things like having multi-site, having redundancy around individual parts of the system to strengthen up those individual components. Then we looked at reducing interconnection between our systems, whether it's bulkheads, circuit breakers, multi-clouds. I didn't go into the detail of it, but in the Sauron application, we actually had ring-fenced bulkheads effectively around the connection pools for each of our downstream services.

The Importance of Learning from Incidents

What happened at Ronan Point was a tragedy. Four people died. It was honestly a miracle. It wasn't more people than that when you see the scale of that. The real tragedy though would be not to learn from these things. The only reason I was able to do this talk is because there was an inquiry. There were reports done into what happened. Journalists pushed to find out what caused this. They wanted to try and move people back into this building after this occurred. It took a lot of fighting to convince them that the issues were endemic. It wasn't a gas explosion that was the problem. It was the fundamental construction of the building. It was only because an architect pushed to have the building brought down and broken up piece by piece that we found out all the things that had gone wrong because a normal demolition would have just involved blowing the thing up.

I only got to talk to you about the AWS us-east-1 outage because AWS put out their incident report. You can see incident reports as PR, and of course that's part of it. There is a bit of part of that. This is also how us, we as an industry, learn from each other. I learned so much from reading these incident reports. They're always fun things to read because, also, it's not you, so that feels good. That's always nice. The Cloudflare ones are especially good because they're really self-flagellating in Cloudflare. They beat themselves up probably a bit too much. If they were a person, I'd go and say, "You're all right." They're always good to read because we can learn things. We can live vicariously through them. All these other reports and these articles and videos, they all gave me an opportunity to learn about what progressive collapse was and apply these concepts in my own way in how I think and how I operate.

You're never going to eliminate failure from your systems entirely. You're going to do your best to manage it. Bad things will happen. It's really important that you create the opportunity to learn from that and to also share that learning with your colleagues and, if you can, the wider industry. That's why I started doing conference talks, would have been 20 years ago now. If you want to know a bit more about Ronan Point, this video is a great watch. I'm going to give you a slide download at the end. I wish I'd found that video before I went through all the PDFs because it does explain a lot of what happened in a bit more detail in a nice 12-minute video. These and other resources are available as links from the slide deck. You can download the slide deck here, but it's also available to download via the QCon platform.

Questions and Answers

Participant 1: Very interested to see you talk about how adding capacity to avoid issues can create issues in itself. I'm also wondering what your thoughts are on the timescale to think about adding protections, because when you have a bad incident and you have a retrospective, everybody's inclination is to be like, we need to make sure this never happens again. I'm wondering some of the things you talk about like, have longer timescales or you need to think about them more. I'm curious how you approach advocating for certain kinds of protections with that in mind.

Sam Newman: I don't think in the wake of an incident going, let's make sure this never happens again is a bad thing. I've said this before, but closing the stable door after the horse has bolted is always a good thing. Ideally, we would have closed it before the horse bolted. We can't go back in time. Yes, maybe we should close the stable door. A lot of the time it's like, I've got a stable and a horse. I think that's understandable. I think it's more about prioritization. It's more about when you increase the complexity of a system, that has an implementation cost, that has an ongoing ownership cost. I think it's like any other piece of work. What is it we're trying to fix here? Getting concrete can help. I didn't go into it in detail, but when doing DR planning, for example, it's useful to separate out your restore point objective, which is when you've last got your backup of data.

How much data loss can you qualify? That's going to be a business requirement. What's my window of data loss, or my restore time objective, which is how quickly do I need to start up and running? A lot of the time, those operational requirements are things which you can go and find out about. Then you can use that to prioritize the work. I think as techies, we are often thinking about this world. I think if we put the context of the failure in a business concept or a business context, get a bit more specific about what our system needs to be able to do, and then break down the work and explain how we're going to achieve it, and then it's prioritization. This thing is going to help us make a bigger impact than this thing, and it's going to be less work. Let's do that first.

Then, for me, it just becomes a conversation with the product owner or business owner about, you don't like that, do you? This will fix it, but it's going to cost you £5 million quid. I haven't got £5 million quid. We're agreeing not to do it. Let's move on with our lives. I just think it's that simple. It's not simple, but you get the idea.

Participant 2: There was one sentence that you posted that I haven't been able to strip out of my head, which is, at some point in the building issue, there was the ability for lesser skilled people to build the building faster. It makes me think about the current situation where, with the use of agents and AI, we are able to get lesser skilled people to build software solutions faster.

Sam Newman: We're doing that by getting rid of junior developers and not hiring M&Es. We're doing that from two different angles. It's great, isn't it?

Participant 2: I was wondering, in this context, if you have any thought that could be applied in the current situation, what guardrails we can put in the case of an agent building software?

Sam Newman: Coming back to the analogy of the construction. The two sides to looking at Ronan Point is, was the system that they'd chosen for building the building correct in the first place? The second thing is, did they implement the building based on that system? There are question marks about both ends. I think we can say the same thing from an AI point of view. The models themselves are getting better. Hopefully what they're producing is better. We're getting a better understanding about how to guide it from a context point of view. The system itself, we're hoping is a better system. We still need the guardrails on the other end, the verification. We had a panel from the cohort training. I think there's a bit of agreement amongst us in the panel, is that we need QAs again. The whole move towards shift left around testing had this problem that we suddenly said, we don't need QAs because developers are going to do all that.

No, proper QAs who can think about system properties and verify if the system is working correctly. Like if I had advice for a junior developer right now who can't get a job, it's like, become a QA for this world. A technical QA that can verify the outputs of these systems or become an SRE. Becoming an SRE is very difficult if you've already got a job. I think we need that. I think we need those skills and disciplines back. We need to find new ways of understanding and verifying our software and making it correct. If you want to run an agent swarm that's going to run two weeks and rebuild a bad browser, that's a well-defined problem space. If we're trying to do that for our own software, we need to get a lot better about saying what good looks like. I don't think we're nowhere near it yet. We're getting there.

See more presentations with transcripts