Ep 19: Lessons from CrowdStrike: Managing Risks in IT and OT Environments | PrOTect IT All
HomeEpisodes › Episode 19
Episode 19
Episode 19 Solo

Lessons from CrowdStrike: Managing Risks in IT and OT Environments

Jul 29, 2024 00:15:43
OT SecurityRisk ManagementPen TestingRansomwareLeadership

Watch This Episode

In Episode 19 of "Protect It All," titled "Lessons from CrowdStrike: Managing Risks in IT and OT Environments," Host Aaron Crow gets into the recent CrowdStrike Falcon platform incident that caused widespread system crashes and blue screens of death on Windows machines. Drawing from his extensive IT and OT experience, Aaron explains that the issue stemmed from a routine update error, not a cybersecurity attack. He explores why it had such a significant impact on major entities like airlines and airports.

 

Aaron highlights the critical differences between IT and OT risk management, emphasizing the importance of automated updates, real-time threat detection, and thorough update testing. He discusses the need for comprehensive risk assessment and the implementation of cyberinformed engineering practices to prevent similar issues in the future.

 

Listeners will gain key insights into balancing cybersecurity measures with system reliability and availability and actionable recommendations for strengthening their IT and OT environments.



Connect With Aaron Crow:

 

Learn more about PrOTect IT All:

 

To be a guest or suggest a guest/episode, please email us at [email protected]

 

Read the full transcript

Aaron Crow (0:1.260): What's everyone? What's up everyone? Wanted to dive into this CrowdStrike incident that happened. Everybody's talked about it. Everybody. If you haven't heard about it, then you probably been under a rock, but little bit about me. Eric Crow. I spent a lot of time kind of grew up in my career, started out doing like desktop support and and worked in infrastructure building, you know, networks and supporting.

Aaron Crow (0:29.654): active directory and exchange. So over my career, I've kind of been in a lot of spaces, including even in my OT space. When I first got into OT and working in critical infrastructure, a lot of that time was spent rolling out basic services that we've had in IT for a while. And what do mean by that? I mean, know, antivirus and patching and firewalls and network architecture and things like that. Things that we'd had in the IT world for decades.

Aaron Crow (0:58.674): But really just rolling those out into these O .T. spaces because as we started bringing that technology into these spaces, we started having some of those problems. We had viruses because we were using commercially off the shelf products like Windows and VMware and Cisco and all those things. you know, before it was proprietary systems that they weren't updated, but they also didn't have the same vulnerabilities because they just weren't as much. Right. If you look at a lot of the vulnerabilities that come out today, it's it's it's because

Aaron Crow (1:26.850): you know, attackers go after the most prevalent systems, Windows and all those types of things. So those are the ones that have the issues that we see the most, right? We see malware on Windows machines. We see viruses on Windows machines. Not that Mac or others or Linux are immune to it, but it's because the most prevalent, the most devices are Windows machines. So that's even the case in the OT world.

Aaron Crow (1:56.478): as we saw in this. Now, this CrowdStrike issue, obviously, well, maybe not obviously, just for everybody to be clear, this was not a cybersecurity issue. CrowdStrike is a cybersecurity tool, but this was not an incident from a bad actor that was attacking or a nation state. It wasn't malicious in any way. It was really just an update issue. So what happened on July

Aaron Crow (2:24.430): CrowdStrike released a routine sensor configuration update for its Falcon platform on Windows systems. This update was meant to enhance the security against specific cyber threats and contained a logic error that caused a system crash leading to a blue screen at death cycle. That cycle would boot up, go into the blue screen and then sit there. Why was the issue so widespread?

Aaron Crow (2:54.306): Ultimately, it really only impacted Windows machines. And the biggest reason there was an, or the widespreadness of it was the response, right? So the problem became widespread due to a few factors. Let's just dive into them, right? So automated updates. I've got my tool, it automatically updates. The faulty update was pushed automatically to systems running the Falcon sensor. And all those companies that were using CrowdStrike for endpoint security have typically

Aaron Crow (3:23.886): you're going to auto update. want your latest antivirus updates. You want all those types of things there. So you're constantly getting those updates and you're securing against vulnerabilities that are coming up. But that also meant that the flaw was quickly distributed across their enterprise or across their environments. Global use of CrowdStrike. CrowdStrike is a major player and for good reason. They're a great product. The CrowdStrike Falcon platform is widely used. We have over 24 ,000 customers, including many Fortune 500

Aaron Crow (3:54.683): As a result, this update impacted a significant number of critical systems. saw major airlines and airports. You look at Twitter and there's tons of examples of pictures of people walking and seeing blue screens and think even Times Square had them. So why do organizations use CrowdStrike? CrowdStrike is a leading cybersecurity firm. It's known for its advanced threat detection and response capabilities.

Aaron Crow (4:24.088): We use products like that for a few reasons, right? Is real -time threat detection. They provide real -time monitoring response to cyber threats. So as devices and systems and vulnerabilities and threats come up, they're using endpoint protection. have artificial intelligence and machine learning to detect those anomalous activities and do something

Aaron Crow (4:49.470): They have comprehensive protection from malware detection, endpoint protection, threat intelligence, protecting system data and operation. Ultimately, you want those systems to be operating as we saw with those blue screens and why it was such an impact. Then CrowdStrike has built trust, major corporations, government entities, because it's effective. It's not that a product is infallible.

Aaron Crow (5:18.306): as we see, and I'll dive into more around the impact, but CrowdStrike is popular because it works, right? It's popular because it's, it's, it's been a, a staple for, for organizations. but why was this issue so impactful? that that's really the key here, right? And it's really understanding, all of these details on what, what can we learn from this to, make sure that we,

Aaron Crow (5:44.770): we have a different or we don't go down the same road again. So critical system failures, the blue screen of death, it was by an update led, that update is what pushed those failures, right? It only pushed it on Windows machines and they were basically inoperable. Like they were until a person put hands on a keyboard and fixed the problem, there was nothing you could do. So they would just sit there at the blue screen of death until you do it.

Aaron Crow (6:12.546): You know, if you think about airline check -ins and banking services, hospitals, all of those types of devices, anytime that any of those were hit, you look at airports and there's hundreds of them. So some person has to put hands on and there are, there were some automated, but for the most part, a lot of those systems are segmented for a reason. So it makes it difficult to get to them. Maybe, you know, the ones in the airport or a kiosk, the machine may be on the back. Maybe they have to, I saw pictures of, you know, workers with, you know, on,

Aaron Crow (6:43.879): ladders trying to get to the PC to be able to do the work. It was a fairly complex remediation. So the fix included, you know, requiring technician to really manually boot systems into safe mode or recovery mode, delete that problematic file, which is time consuming. Again, just that process, booting it in, even if I can do it remotely is difficult. But when you have other things

Aaron Crow (7:12.022): bit locker and the fact that these devices are physically segmented or I have to go physically put hands on them and there's tens, hundreds of these devices spread across large geographical areas, every airport, all these different locations. I don't have necessarily people just sitting there waiting to deploy. So how do I put hands on all of these things? And then the simultaneous failure systems worldwide.

Aaron Crow (7:39.074): it really compounded on each other, right? So we had this continuity issues, large companies, airports, you know, if you look at the airlines in the air during that time, there was a significantly less amount of airplanes in the air and that was compounded because this was not necessarily an OT issue or obviously it wasn't a cybersecurity issue, but this is a prime example of how this IT and OT convergence thing, right? So we have these systems and many of these systems were OT systems.

Aaron Crow (8:8.078): in my view, but they weren't the systems on the on the airplane. Like they weren't not flying planes because there was crowdstrike on the control system of the plane. It was because they couldn't book things that couldn't schedule. They couldn't get you boarding passes. They couldn't, you know, book your your your your luggage. Like all of those were the reasons that brought this thing down. And that that jaws into the bigger the bigger picture of why things are different in IT and OT.

Aaron Crow (8:38.510): automatic updates. There's a reason why we don't patch in OT the same way we do in IT. It's a reason why, you know, one of the other stories in this is Southwest Airlines came out and they had Windows 3 .1 and they weren't impacted the same way that some of the other airlines were. I'm not saying that everybody should have really old operating systems, but what I am saying is, is it's a prime example of upgrading and having the latest and

Aaron Crow (9:7.808): of everything doesn't necessarily make you more reliable, make you more available or make you even more secure. At the end of the day, these entities are doing what they need and what they focus on, whether it's a power plant, an airline, an airport, a train, a manufacturing facility, just replacing and upgrading is not necessarily the right choice and not the right action. Same thing with with updating, right?

Aaron Crow (9:35.628): We want to update our CrowdStrike or whatever our systems are, the iOS on our Cisco devices, the firmware on our firewalls, all these things. In an IT world, I'm going to update those almost instantly or very quickly within weeks of those things and sometimes hours of those things being released. But in an OT world, it's so dangerous to do that. And this is a prime example of how that is dangerous. that also brings in risk.

Aaron Crow (10:4.768): If I update these devices, then there's a risk of bringing down an entire entity. We saw this on a large scale. This may be one of the largest ever, but what do I do about that? If I'm not going to, let's play devil's advocate and let's pretend that we're not going to auto update any of our devices and critical systems from now on. What does that mean? Well, how can you do this differently? Well, you can roll it out to a smaller group. You can have a test environment.

Aaron Crow (10:34.162): You can validate that it's not going to break because if they'd tested this on a few systems before they just broadly rolled it out across their entire organization, then they would have seen these blue screens come up and they would have stopped it from impacting their entire operation. But that takes time. That takes resources. That takes dedication. How many updates are made? How frequently can they do that? And if they're not updating on that same schedule, I mean, some of these updates happen daily, hourly.

Aaron Crow (11:1.034): And then what happens when you don't update it and that vulnerabilities there and you've got a hole and a vulnerability in these environments. So you have to look in architect these environments, purpose built. talk a lot about the cyber informed engineering. Idaho national labs has done a great job of, really pushing that concept and designing this idea. I'm going to be speaking about it at, at DEF CON and ICS village, but ultimately it's, it's around

Aaron Crow (11:31.192): Purposes like this, right? What can I do? Let's take a step back again, going back to, can't patch these things all the time. So how can I make sure these environments are safe and secure and available when I know I can't patch them all the time? So I have to make other remediations and other mitigations to fix those things. I need to make sure I'm having backups. I need to make sure that I'm doing testing. I'm going through all of these steps and I have somebody at the table that's playing that devil's advocate. Because obviously it makes sense to go put patch and put the latest,

Aaron Crow (12:0.684): vulnerability information for my endpoint protection, like a crowd strike, right? Obviously, I want to have the latest and greatest. I want to have the latest information. It's like the president is always wanting the latest information about whatever is going on in the world. He doesn't want to be working on, you know, two, two month old or even our old information because it could, it could have changed and his decisions can, can vary depending on, on that information. We look at it the same way, but obviously on the flip side, the risk is something like this can happen. so we've got to have.

Aaron Crow (12:30.348): You know, there's a reason why OT is segmented. There's a reason why we don't pass at the same level. There's a reason why we use different tools. There's a reason why we segment from, you know, Active Directory and we don't put OT devices typically into an IT organization. We have different teams that are pushing it. Like you need to have different analysts that are looking at the data and all these things are because a lot of the technology that's in OT and IT are very similar, but the way that we run them and the impacts to a down.

Aaron Crow (12:59.854): an outage. If your email server goes down for a few hours, it's a bad day, but it's not the end of the world. Your entire organization doesn't shut down. But as we see with incidents like this, an OT environment bringing down your entire, you know, booking operation means that you can't fly planes. It means that, you know, you can't, you can't sell tickets. You can't, you can't, all of these things have these ripple effects across sectors. So, you know, the crowd

Aaron Crow (13:29.314): Blue screen of death, this incident, you know, highlights the critical dependency of organizations and how critical it is on cybersecurity solutions, right? And the cascading effects of how a single update can have on global operations. Understanding the cause and the widespread impact of this issue will help to really underscore the importance of how update testing and challenging maintaining cybersecurity.

Aaron Crow (13:58.292): while CouchTrack has taken all the great steps to rectify the problem, it really just shows it's not a CrowdStrike problem, right? This is yes, this incident came from CrowdStrike, but it's, it's not, it's an, it's an industry issue, right? It's a, it's a, I deal with my OT environment, how I understand risk. it's really a learning experience or it can be and should be a learning experience for both cyber providers, vendors, clients, asset owners to really understand the importance of the comprehensive management of our environments and architecture.

Aaron Crow (14:28.478): and how to, know, what my recovery plan is and how do I make sure I minimize these impacts in the future. Dig into CrowdStrike, they had a good response and detailed information about what happened. You can look at CrowdStrike's page, Wikipedia, and a lot of folks out there really talking about the details, the technical nitty gritty of what happened, but from a business and an OT and an overall just, you

Aaron Crow (14:58.340): architecture side, there's a lot to be learned from this

Transcript lightly edited for readability.

Want your brand in front of OT, IT, AI, and cloud security decision-makers?
PrOTect IT All listeners are the practitioners and leaders making security buying decisions across critical infrastructure.
See Sponsorship Packages →

Never Miss an Episode

Subscribe to PrOTect IT All and stay ahead of the threats targeting critical infrastructure.