In this episode, Aaron is joined by Paul Shaver, an experienced OT security consultant from Mandiant, part of Google Cloud. Together, they navigate the nuanced landscape of operational technology (OT) cybersecurity.
The episode begins with Aaron recalling a critical incident at a power plant that underscores the potential pitfalls in OT environments. This sets the stage for a rich discussion on the evolution of OT technology, with Aaron and Paul reminiscing about primary domain controllers and early NT workstations.
The conversation shifts to the future of OT in the cloud, where Paul highlights the benefits of cloud solutions, including enhanced resiliency, security, and data optimization through AI. A compelling customer case study illustrates modern technology adoption with web-based HMIs and Chromeboxes.
Paul offers a detailed analysis of the current OT cybersecurity landscape, addressing the persistent legacy system challenges and the need for a cohesive IT-OT security strategy. He discusses the evolving threat landscape influenced by global geopolitical tensions and the rise of zero-day vulnerabilities.
Listeners will gain practical insights into foundational cybersecurity measures, such as network segmentation, asset inventory management, and robust access control..
Key Moments:
04:14 Connecting IT and OT optimizes processes securely.
09:54 Lost production severely impacts manufacturing revenue recovery.
14:06 Ensure network notifications; control access, separate credentials.
17:10 Engineers need secure access to adjust parameters.
21:55 Endpoint detection on older systems is critical.
28:47 Resilience is crucial in CrowdStrike incident response effectiveness.
32:11 Limited resources for global incident response efforts.=
39:22 Rebuilt domain controller caused authentication issues.
42:37 Focus on resiliency and cloud opportunities, leveraging multi-cloud.
44:59 Improve grid operations using cloud and hyper-converged technology.
48:38 Local cloud provides redundancy for remote sites.
51:15 Critical for acquisition process and problem-solving.
About the guest :
Paul Shaver has dedicated more than two decades to various roles in Operational Technology (OT), primarily within the oil and gas industry. His expertise spans OT architecture, design, and build, along with run and maintaining responsibilities as an asset owner.
Before transitioning into cybersecurity, Paul served as a Technology Director for an oil and gas company in California. Driven by a burgeoning interest in security, he joined Mandiant nearly five years ago. At Mandiant, now part of Google, Paul relishes the mission of enhancing security postures in OT and critical infrastructure, contributing to significant advancements in the field.
How to connect Paul: https://www.linkedin.com/in/pbshaver/
Connect With Aaron Crow:
Learn more about PrOTect IT All:
To be a guest or suggest a guest/episode, please email us at [email protected]
Aaron Crow (0:1.174): Welcome to the show, Paul. Thank you so much for taking time. excited to have this conversation. We've, we've met many times in person and, we've, we've crossed a lot of the same paths. Obviously this, this world is fairly small, especially in OT cybersecurity. So why don't you introduce yourself, tell the audits, who you are, what you do and, who you work for.
Paul Shaver (0:20.394): Yeah, awesome. Thanks, Aaron. Thanks for having me. So Paul Shaver, I work for Mandiant, which is now part of Google Cloud and Google Cloud Security. I lead the ICS and OT Security Consulting Team at Mandiant. And I've been here about five years. Prior to that, I've spent about 20 to 25 years, various roles in OT.
Paul Shaver (0:45.876): Spent the bulk of my career in the oil and gas industry and the majority of my career in OT architecture, design, build, some time in there with run and maintain at an asset owner. Prior to joining Mandiant, was a technology director for an oil and gas company in California and decided I wanted to make the full-time jump to security. Had spent a lot of years with security kind of as a collateral role doing those jobs. And I'd landed at Mandiant five years ago.
Paul Shaver (1:15.284): just shy of five years. And yeah, it's been a fun experience and really enjoy the mission and what we get to do here. And now being part of Google, really expanding that and helping to improve security posture in OT and critical infrastructure and all of those buzzwords, the world that we live in.
Aaron Crow (1:37.076): Absolutely. Yeah. And, I definitely want to get there, but first I just want to kind of kick us off with, you know, kind of current state of OT cyber, you know, from, from persistent changes to, you know, legacy systems, visibility, like all of those basic types of things of the current state that we're at. Like you said, like we've both been doing this for 20 something years and some of the problems that we're having, we've been, we were having 10 years ago and we were having
Aaron Crow (2:4.494): 20 years ago and some of the same systems are running that were around 20 years ago when we started this thing, right? So kick us off with where we are today from your perspective and kind of this OT cyber space and where we sit.
Paul Shaver (2:5.343): Mm-hmm.
Paul Shaver (2:20.638): I you know, we still, like you said, to your point, we're still facing a lot of those same struggles, right? The same things that add complexity, the same things that introduce risk, right? That expand the attack surface for these environments, kind of all still there, right? We still have legacy systems. We still have a lot of Windows XP running HMIs in these environments, right? So there's a lot there. I think what we're seeing probably last two or three years,
Paul Shaver (2:50.970): More and more companies are starting to take the OT security of their estates, of their environments more seriously. And I think we talked about this a couple times when we've hung out. We're starting to get to a point where we're not chasing a bunch of legacy problems and we're getting the chance to innovate a little bit. The technologies are improving, the vendors, the...
Paul Shaver (3:18.482): in the OT device space are starting to build security into those OT devices, the PLCs and the communication cards. so now we're starting to be on this kind of level playing field where we see green fields being built and we've got a bunch of really good security capability that we've never had before. And so I think the life cycle on security is starting to get better with new products.
Paul Shaver (3:46.282): We're starting to catch up with this idea that, you know, we can't do this. We can't implement this, you know, this, control in this environment because the PLC doesn't support it or the com car doesn't support it. so there's, they're starting to balance out, and, now, you know, again, the, the, also the expansion of the technology that's connecting to these OT environments, right? Organizations are connecting their ERP systems, their manufacturing engineering systems.
Paul Shaver (4:16.074): they're using data to optimize their processes or they're using data from the production floor to optimize financial process, right? So with those connections, we're seeing companies start to look at how they can connect between IT and OT and leverage that data. And so that brings in both a complication to how we do that and how we do that securely, but it also brings in the opportunity to have some really good conversation about
Paul Shaver (4:45.738): you know, what does data flow inside the OT environment look like? Not just what are we connecting it to in the IT side? You know, how can we better segment? How can we improve efficiencies in just network efficiencies and bandwidth and, and latency and all those things in the OT environments, because we're starting to have conversations around security and introducing external connections. So I think the security conversation facilitates a lot of like really, how do we better
Paul Shaver (5:14.582): How do we optimize these environments in ways that then we can leverage for security purposes as well? So a lot to unpack in that statement, right? I do see like five years ago, six years ago, would have said, everything's a mess. Everything's vulnerable. Like we got to fix this. And I think we're getting to a point now where we're seeing more companies spend the money.
Paul Shaver (5:43.010): invest the resources, people, time to make those improvements. And that says a lot for how fast we're addressing some of these problems.
Aaron Crow (5:54.196): It seems slow, but it is really fast, right? Especially in these spaces. So pivoting off of that and continue down that thought process, have you seen the threat landscape shift in the last three to five years? We've seen so much more in the news around critical infrastructure and OT and stuff like that. So what have you seen in the threat landscape focusing on these environments in these OT spaces? Because we just see more and more in the news.
Paul Shaver (6:25.066): Yeah, I mean, a lot of that has to do, I think, with kind of geopolitical struggles across the globe. And so, you know, where we see threat actors from, you know, Russian, Russian associated groups, China associated groups, you know, trying to establish a foothold and some kind of persistence in these environments, leveraging edge devices and the vulnerabilities that, know, I can't remember the number. just...
Paul Shaver (6:54.730): I just had this stat for another call, but a ridiculous amount of zero days in the last 12 to 18 months. And a lot of those on edge devices, firewalls and load balancers and some kind of edge connectivity. for OT environments, means a lot of those, maybe it's a cellular motor, maybe it's a firewall that connects to some kind of wireless backhaul.
Paul Shaver (7:24.264): or some kind of connection that leverage a public infrastructure, that means you've got something potentially hanging on the internet that's vulnerable. That's a persistent point. That potentially becomes an endpoint for use in a botnet. There's a lot of that. So I think because we've got these threat actors that have some geopolitical type stake for leveraging that, that...
Paul Shaver (7:52.564): the increase in vulnerabilities and zero days that we've seen exploited. you know, there's a lot more reason to be concerned, but I also believe that we're doing a much better job of, improving the security posture of these environments. If you go back, you know, that five, six years and you run a show Dan scan for, you know, Rockwell PLCs or HMI devices or whatever it might be, you know, there were a lot of stuff connected directly to the internet.
Aaron Crow (8:20.951): a lot.
Paul Shaver (8:22.640): And now when we run those research scans, when we're doing those type of engagements for clients where we're looking to see what their attack surface that's on the internet looks like, it's way better than it used to be. So yes, the threat landscape changes, the attack surface broadens with some of these zero days and more edge devices. But I think we are doing a better job of identifying that stuff early.
Paul Shaver (8:49.626): The threat intel is looking for these things and doing a better job of getting victim notifications out there. So we're definitely evolving on both sides of the coin.
Aaron Crow (9:2.414): Yeah, absolutely. It's changing so much and it's just a, it's an uphill battle. and it's constantly changing cyber is in general, but you know, OTA, I think we all believe we're a little bit behind the eight ball, but there's, reasons behind that. We have to move slower because there's, there's physical implications to those things. Like we can't just go patch everything and hope for the best. you can't do that in a production environment because it'll, it could kill it and literally kill people.
Paul Shaver (9:23.934): No. Yeah.
Paul Shaver (9:28.532): Yeah, mean, worst, you know, the best case is, you know, some kind of a equipment damage in those cases, right? Then you move to like some kind of environmental damage or loss of life or limb. there's, you know, and you take all of those those risks out of it. And it's just a revenue risk. Right. These companies that are running a production environment, manufacturing a widget, you know, three days of lost production is
Paul Shaver (9:58.300): You that's three days of lost revenue potentially. Right. And you look at some applications in like refining or chemical processes or oil and gas, you have to shut in production for two or three days. might take you weeks to get back up to that same level of production because you have to get stuff up to an operating temperature or you have to, you you have to. You you've got, you know, stuff being made in a, in a batch sequence. And when you shut that process down,
Aaron Crow (10:0.472): Correct.
Aaron Crow (10:25.709): Right.
Paul Shaver (10:27.722): You have to start back at batch zero and that might take four or five, six days a week or plus to get back to that optimal level of production. So, you know, there's definitely a revenue impact to being able to patch and update in some of this. So you really have to look at compensating controls and good security hygiene and make sure that we're doing the things that are most critical to protect these environments.
Aaron Crow (10:29.518): All right.
Paul Shaver (10:57.844): You do all that before you start spending millions of dollars on network security products that are great, that do a really good job of helping you identify the things that are in your environment. But if you don't have good architecture, or if you don't have good segmentation, if you don't have good access controls, then your network security monitoring maybe is just going to detect a thousand things a day because you've got lots of holes in the sieve, as it were.
Aaron Crow (11:7.056): Sure.
Aaron Crow (11:25.614): Which leads me to my next topic and it's really around when it comes to foundational cybersecurity hygiene, like what are the must-haves for an organization to protect their OT environment, right? You just kind of talked about that at a brief level, but it's not all spending billions or millions or hundreds of millions or huge numbers. Sometimes the foundational things that we need to do.
Aaron Crow (11:49.494): are just basic, like they're processed, they're people. They're not always buying the new technology and the Lamborghini. Sometimes you just need to buy the screwdriver.
Paul Shaver (11:58.602): Yeah, no, absolutely. The biggest one is network segmentation. can't tell you, and I'm sure you know this, that you work in this space and you go out and support customers much the same way our teams do. The large flat networks in these OT environments still exist. The application servers, the engineering workstations, the HMI,
Paul Shaver (12:28.630): machines in the control room and the PLCs and the comm devices, sometimes even the network connected instrumentation, all still on that same subnet, all still on one default VLAN one or VLAN 255, and all addressed in a private class C address space. And good Lord, there's so much out there. So network segmentation is the first thing.
Paul Shaver (12:57.706): in kind of like peeling back the layers of the onion to figure out what is on these networks. The number of devices that you have in asset inventory is kind of the next thing and trying to understand what's there so that you understand what the risks are. If you've got HMIs that are running embedded Windows XP, that's one level of risk.
Paul Shaver (13:26.574): versus maybe those are, those are a very specific Linux kernel that was developed just for that machine. And, and there's no, you know, there's, there's a list of CVS for that. It's like, you know, two, three CVS for that device versus I, I don't even remember what the list of CVS for windows XP is. It's in the, it's in the thousands, if not tens of thousands. Right. So, understanding what you have becomes really critical.
Paul Shaver (13:55.988): These are again, all basic security hygiene stuff and then controlling the access into the environment, making sure that the people that, you know, the machines, you know, we can't necessarily zero trust a, an OT environment. Again, the technology doesn't support it, but you can ensure that, you know, if a new machine connects to the network, at least somewhere there's a, there's a, you know, notification from the switch, from the router, from the firewall, from the
Paul Shaver (14:25.440): from the infrastructure somewhere that says, hey, something new connected. We need to know what that is. having an idea of what's on your network is important, and then controlling who can connect to it, who can access those environments. Not having shared credentials between your IT and your OT environment is, I can't tell you how many times we see, even where
Paul Shaver (14:54.090): you're not enforcing good password security in the OT environment. So the username is maybe different between IT and OT, but that guy that has access in both environments, he doesn't want to remember two passwords, and so he's using the same one on both sides. And that's on a sticky note under his keyboard.
Aaron Crow (15:16.099): are printed on the monitor that is sitting in the room.
Paul Shaver (15:20.106): Yeah, the admin password is on a label tape on the side of the HMI out in the field because the operators need the admin password when they have to change a set point.
Aaron Crow (15:29.954): Exactly.
Aaron Crow (15:35.062): Well, and that's, that's, gets to a good point too. And I've had this argument conversation with people, especially people outside of OT that don't get it. Like you work into, you walk into a control room and a power plant, they're, not the workstations are not locked. They never have to log in with a password, but there's a reason for that, right? Is, a, first of all, there there's multiple physical layers of security that you had to get to B there's not that many people that work there. And if you're not supposed to be sitting there,
Paul Shaver (15:53.141): Mm-hmm.
Aaron Crow (16:4.512): somebody's going to notice it. It's, man, 24 hours a day, seven days a week. They've got cameras on it. Like there's a lot of other mitigating factors that go into that. but at the end of the day, if there's something going on and they need to stop something like, you know, life liberty, you know, all this stuff, like you can damage equipment, you can kill people, like you can blow up, you can all sorts of really bad things can happen. Any hesitation from that operator being able to stop or start a process.
Paul Shaver (16:24.832): Mm-hmm.
Aaron Crow (16:32.322): that can save a life or all that is more important than having to log in with a password and forgetting his password and having to type it three times. Because we've all done that. You have to, I know my password, but I have to type it six times because I fat fingered it because it's a freaking complex.
Paul Shaver (16:34.571): Yeah.
Paul Shaver (16:41.430): 32 characters. Yeah.
Paul Shaver (16:48.614): Yeah, 32, 32 random characters and, know, 10 % of them have to be special characters and, you know, four capitals. Yeah, it's, no, it's, it's, it's a completely valid point, right? We don't, the, the operators don't need, you know, shouldn't need to, to log in, right? You typically, especially when you've got something that's man 24 seven, right? There's that person's always there.
Aaron Crow (17:15.992): Sure. Yeah.
Paul Shaver (17:18.853): But, that might just be, operators might just have the ability to do set point changes or have a restriction on how much they can adjust that set point. Whereas the engineers that can actually change logic, change code, set that set point outside of a predetermined set of values, that absolutely needs a password. That absolutely needs two factor authentication if you can enable it.
Aaron Crow (17:41.710): 100%.
Paul Shaver (17:47.506): and not having those shared credentials. And then again, so it's policy too, right? It's having good security policy that is collaborative with the IT security policy and enforceable in the OT environment, right? You don't want these two drastically different policies because that becomes a nightmare to try to manage it. But you also...
Paul Shaver (18:13.610): If you've got an HOT environment and you can't use zero trust or you can't use two factor authentication or your endpoint detection tools agents can't be installed on those machines, your policy has to be able to be collaborative with your IT environment to be able to make sure that it's manageable.
Aaron Crow (18:36.290): Yeah. Yeah. I mean, I have a great use case of that is, is IT was pushing policies down in an OT space. and they were, they were doing real good work and they were, they rolled out a group policy through active directory that just locked workstations after five minutes of inactivity or 15 minutes, whatever the number is. But there was literally a screen, it was on the corporate network. So it wasn't technically an OT device. Like we have this argument all the time about, this an IT attack? Was it an OT attack?
Aaron Crow (19:4.694): At the end of the day, just like calling a pipeline, none of that was OT, but it impacted OT, right? They shut down the pipeline because of it. So does it matter if it was an OT attack? No, it really doesn't because it shut it down, right? So same thing happened in this scenario. This was in a control room. It was a screen that had PI data. And as we know, the operators run off a PI, right? Everything they do, they're running from PI and all the indications that come from that OSI PI server. Well, they locked the screen. They didn't know what the password was because the operators never had to log into it.
Paul Shaver (19:11.413): Mm-hmm.
Paul Shaver (19:24.075): Mm-hmm.
Aaron Crow (19:34.380): they had to get an engineer who was not there. He wasn't on shift that day. So the person that knew how to log in couldn't and they couldn't get the data they needed to because an IT person trying to do a good job and put this device into a group policy that locked down the workstation and they had no access to it. Right. So it's a prime example of how a simple, good policy, good hygiene conversation, like nobody would argue that was a good idea except that
Aaron Crow (20:1.698): They didn't understand all the implications that it was going to bring. And obviously it got put into a separate group policy and, that goes back to the whole asset inventory. They knew that device was there, but they didn't understand the function of it. Right. So yes, they knew it existed. Yes, they knew they were patching it. They were backing it up, like all that kind of stuff, but they didn't know what it was really used for. So they didn't know that when I rolled this thing out, there could be implications to my production at a power plant. Right. And God forbid it's at a nuclear power plant that they can't control that. That's a whole nother line of problems.
Paul Shaver (20:30.686): And it is, it just comes down to that good hygiene. So when you're, and I love the fact that there's an OT environment with active directory in it. This is another thing that we just don't see enough of. We don't see access credentials being managed even to that level. So definitely having an OT environment that has its own active directory, phenomenal.
Paul Shaver (20:55.850): But it just comes back down to asset inventory and access controls and knowing what's there and making sure that you're building your Active Directory, you know, forest out, you're creating OUs, you're creating user groups that include these, you know, these machines and being able to categorize them for their purpose. But to your point, you know, 99 % of what we see that impacts OT environment starts in IT, right?
Paul Shaver (21:21.800): It's got to, it has, whether that's an IT system on the enterprise side or that's an IT system on the OT side, it doesn't matter. These attackers have to have for the vast, vast majority of these cases, they have to have that, that access point, that initial compromise. And that's going to be the system that back to that, you know, the system that has the most CVEs that potentially can't get patched.
Paul Shaver (21:51.482): That's going to be a Windows operating system. That's going to be a Linux operating system. That's going to be something that is a traditional OS sitting in that environment that gives an attacker an initial compromise and a place to pivot from and do their recon. protecting those assets becomes the most critical. Again, all the network security monitoring platforms that are out there, they do a phenomenal job of being able to detect anomalies and
Paul Shaver (22:19.390): and help organizations better understand what's in their environment and what's happening in their environment. But that not having endpoint detection capabilities on those operating, those commercial off the shelf operating systems that we're beholden to is critical. And so for some of those older systems where you can't install an endpoint agent or you can't have EDR running or antivirus running, we really need to have that visibility of,
Paul Shaver (22:49.807): audit logs and system logs and application logs and pay attention to what's happening on those machines in a much better way. Again, good hygiene methods here where before you're spending all this money, know what data is available that can be leveraged for a security purpose without necessarily having it be quote unquote security data.
Paul Shaver (23:18.186): PLC system variables are a great example of this. tell people all the time, if you can't afford network security monitoring platform, that's right now, let's get you to a point where you can. But in the meantime, CPU runtimes, PLC scan, ladder logic scan times, memory usage on a PLC, those are all system variables that can be archived in your historian. And just like they're archived in their historian, then
Paul Shaver (23:46.536): Once they're there, they can be used to create an alert in your HMI alarm system. And that can be escalated based on, the CPU runtime is abnormally high. We can say, hey, something changed. Let's go figure out what changed. Can we tie that to a change management process? Look, you have a security triage playbook right there in front of you. And you don't have anything but the normal process data that
Paul Shaver (24:15.040): We're just looking at it in a different way. So that's really valuable for the small operators that can't afford, they don't have big security budgets, they can't afford the tools that, and maybe they're working towards that, but there's lots of ways to start at the ground level to protect these environments and detect when something changes.
Aaron Crow (24:36.534): Yeah. And when I started this process, it was before we really called it OT cyber, right? And a lot of the things that I was rolling out, I didn't have a budget, right? As I was working as an asset owner and a power utility. And, you know, I've told this story multiple times, but you know, when I, when I went to the plants and was explaining what we were doing and that what we're trying to do, I didn't sell them on cyber because again, this was, you know, 2010, nobody cared about cyber. It wasn't a concern that these asset owners are really all that focused on.
Aaron Crow (25:6.638): What they were caring, they did care about is their bottom line availability, you know, anything that could make them run more efficiently. So what you just talked about, right? So helping them monitor. like we were getting logs out of all these systems that they already had and we were just showing it to them in a way they hadn't seen before. Right? So when we first turned on our first Splunk server in like 2010 at a power plant in an OT environment, monitoring a critical, you know, a DCS control system. And one of the very, within seconds we started noticing it bubbled up to the top.
Aaron Crow (25:36.460): We found that there was a redundant switch that they didn't know it was sitting there beeping and having an error code. And it was telling the world, but nobody was listening, that it had a problem. And it was actually, it was down. Like it had administratively taken itself down because it was overheating. Well, we went over to this device, we found the device and we went over to it and the fan wasn't spinning. So it said, you know, it's overheating, fan's not working in the power supply. we walked over to it, it a Cisco switch, walked over to it and there was a zip tie.
Paul Shaver (25:52.181): Mm-hmm.
Aaron Crow (26:6.880): in the fan, like through the case, blocking the fan from spinning. So we pull the zip tie out, the fan starts spinning, the error code goes away. But nobody had ever seen that or walked by it and understood what that red blinky light was or logged into it and known. So these tools, they're not just cyber related, right? They're efficiency related. They're the reliability, the resilience. It's more than just cyber. know we work with the, know, cyber is the big thing right now, but
Aaron Crow (26:35.746): at a power plant if they have to choose, it really comes down to availability or a lot of these critical infrastructures. They care about availability and safety, safety and availability. Cyber is one risk mitigation. It's one attack vector, but they really care about that availability. So if all of these tools, yes, they give cyber data, but they can also make it more reliable and give data to operators that can help them run their plant better.
Paul Shaver (26:57.536): Right. Yeah, no, I mean, and that's, that's the name of this game. Right. And I think we, we, we focus a lot on security and security is important keeping these systems. But I, but I look at it from a resiliency standpoint. Right. And, and to that point, uptime is the most critical factor, right? You're whether you're a power grid, you know, in the hottest part of the summer or you're a gas producer in the winter or you're
Paul Shaver (27:28.179): pumping the wastewater out of the sewer so it doesn't back up in somebody's house. Uptime is critical. And so a lot of what we look at in better protecting these environments is also...
Paul Shaver (27:45.270): putting resiliency at the forefront of this. We want to layer in those levels of defense. We want to be able to detect and respond more quickly so that we keep these systems as resilient as possible and keep up time. But also, that also enables us to have quicker recovery capability. So when we do have a problem, when there is a compromise, when we don't know that if or how,
Paul Shaver (28:15.092): an OT system is compromised when the IT system is the tools that we have in place for, for response capability, for forensic capability, for detection, speed the recovery time. Right. I have the, use the hurricane analogy a lot, right. If you, if you've lived on the Gulf coast and you've dealt with hurricanes, you know, that you're not going to stop a hurricane from making landfall. Right. But if you've got good defensive measures in place, you board your windows up, you sandbag your doors, you
Paul Shaver (28:44.948): You are keeping the trees away from, you know, your, roofline of your house. So they're not going to blow on it. you're building good defense for capability. and then, you know, the resiliency is there with recovery time shortens and you have less overall impact from a, from a storm. then you would, if you didn't do those things. And so the cyber defense comes into the same kind of an aspect of better prepared means quicker recovery, better resiliency.
Aaron Crow (29:14.028): Well, and we had a great example of this worldwide incident with CrowdStrike, right? And we saw a difference and really the difference in how people responded were how they had built resilience into their system. Like how well did they have they tested their incident response and recovery plan? How well have they tested their backup plan? How well have they gone through all these steps? Because
Aaron Crow (29:39.295): At the end of the day, yes, it was a patch and it was a Microsoft thing and all the different key factors that went into it. But ultimately it was the owner of the systems in my perspective of this is just speaking by Aaron, not by anybody else, but it was ultimately their fault, right? You I would never have pushed updates to my entire fleet of systems in a power industry when I supported 45 power plants.
Aaron Crow (30:6.870): I would never have just blindly sent them all to one. Like I had a staged way that I did that. I'd test one system, make sure it didn't break anything, and then I'd roll it out to another system. Okay, it didn't break the second system. And then I'd start rolling out to more people, but I would do that per site. I didn't just blindly say, it worked okay in the lab, send it, right?
Paul Shaver (30:19.766): Mm-hmm.
Paul Shaver (30:28.554): At five o'clock on a Friday. And, and the, to your point, like the recovery time. is it, I had a, I had a meeting with a, with a customer that week. I think it was probably a day or a day and a half later. and I got on the call a couple of minutes early and I was fully expecting him to not show up because I know he's a crowd strike shop. I, I was like, I hadn't heard from him.
Paul Shaver (30:55.568): And so I sat on the call and when he, when he popped on, was like, look, I said, if you want to reschedule this, I'm sure you're busy doing this. He goes, that we're fine. He goes less, less than 18 hours. He said it's like 16 hours and 27 minutes or something. Everything was back up and running. And I was like, wow, how'd you do that? And he goes, disaster recovery systems were tested and they worked and we had a plan and a process in place. He said, admittedly, he said there was a guy that couldn't sleep and he was up at one o'clock in the morning and saw something weird happen.
Paul Shaver (31:25.410): And we got on it probably quicker than most people were able to get on it. He said, but even if we hadn't had that, he said, we'd still be talking about 24 hours rather than 18 hours or whatever it was. so that having that plan in place, having that capability and knowing that it's tested and, and, you know, you can push a button and roll it back, is, incredible. And, know, if you know about the crowd strike problem, right, that's not just, that's not just a
Paul Shaver (31:54.626): roll back the application. That was a boot level process that needed to happen and their recovery processes supported that, which is incredible. that didn't have a, it had a minimal impact to their production environments, but it did have an impact to their production environments because they did have application servers and data servers that had the endpoint agent on them.
Aaron Crow (32:4.802): Yeah. Yeah.
Aaron Crow (32:23.074): Yeah. Yeah, it's insane. I have a couple of stories from folks. Same thing having having those conversations and it was either people were down for a week and they were panicking because, you know, Delta Airlines is a prime example of that where they had kiosk all over around the world and it wasn't a problem if they couldn't do it, but they had to physically touch things. So they it was it was a it was a body problem. So, you know, we sent people out. They were getting people from every consultant firm they could find. And that gets to a bigger problem of, you know, making sure that,
Paul Shaver (32:24.576): Be prepared.
Aaron Crow (32:52.982): a lot of folks will put their, their, all of their incident response plan into a, we've got a retainer with Mandiant or whomever, right? And, and as good as you guys are, and all of us are, there's only so many of us, right? And when a global incident like this happens, there's only so many competent and qualified people that we can put on a plane and send out to put hands on keyboard. So there, there's a problem, a resource problem from that perspective, because everybody uses Mandiant because you guys are
Paul Shaver (33:15.573): Yeah.
Aaron Crow (33:22.654): lead, you know, industry leading in that space. Rightfully so. But again, how many people do you have that you can send out to do this? And when every company, when you talk to Amazon and Facebook and Delta and American Airlines and Southwest or well, Southwest didn't get it because they were running really old systems. But that's a whole nother conversation. But that. Yeah. Yeah, but you know that that's a bigger problem in part, and it goes to the.
Paul Shaver (33:40.877): Yeah, yeah, I your systems don't even support endpoint agent. Back to the XP, right? Yeah.
Aaron Crow (33:51.362): the overall resilience and when we talk about resilience, we're not just talking about technology. We're not just talking about architecture. Yes, that is a factor and that is a piece of building a resilient system, but it's also the, the, you know, bright glass, you know, the, shit handlebar, right? It's like, we know something's gonna happen. Have we planned for that? Like, have we thought through some of those things? And you're never gonna think through everything.
Aaron Crow (34:16.024): But if you go through that, that exercise thinking, okay, I know I designed an awesome system. I'm really smart. Good job, Aaron. But what could go wrong and bringing in outside people? Have you thought about this? Have you thought about that? Like going through those exercises. And I think that's the biggest piece of people that, probably came out a little better than others were because they probably had those conversations and their system was more resilient because of that.
Paul Shaver (34:38.942): Right. Yeah. and, and you're absolutely right. So incident response retainers, right? Yes. You should absolutely have an incident response retainer with, with Mandiant. we've got a free one by the way. So, I mean, that just gets paperwork in place, you know, worst case scenario, we, it takes, we don't go through the, however long it takes to get contracts signed. You should have that capability with as many of
Aaron Crow (34:48.814): Sure. Yep.
Paul Shaver (35:6.496): providers as you can have because when,
Paul Shaver (35:11.060): When a bad day happens, it's not always possible to answer the phone. And to that point, when you've got this mass conflag situation, if we had this huge compromise that was affecting the entire globe, it's going to be all hands on deck. And Mandiant and CrowdStrike and Palo and all of these organizations that have incident response teams are all going to be working together to solve it. We've seen that happen before. We know that the collaboration is there when the
Paul Shaver (35:39.754): We're all in competition on a day-to-day basis in some cases, but when those bad days happen, this is a critical mission field and everybody comes together and we get the job done. But having contracts in place and not having to wait that additional 24 or 48 hours to get paperwork signed, or especially during a compromise when you probably are going to want that contract to be a three-way contract with counsel and you know,
Paul Shaver (36:9.556): That's going to make that process a little bit longer. Council is not going to want to accept web terms or whatever the base level terms and conditions are. So go through that exercise. I can't stress this enough. Get the paperwork done ahead of time. I know that it's, in some cases, you're maybe spending a little bit of money, but having that level of assurance that somebody is going to answer the phone is well worth it.
Aaron Crow (36:40.044): Hang on just a second.
Transcript lightly edited for readability.
Subscribe to PrOTect IT All and stay ahead of the threats targeting critical infrastructure.