Corelight Bright Ideas Blog: NDR & Threat Hunting Blog

The Water-System Attacks Were Simple. Securing OT Isn't. | Corelight

Written by Paul Mathis, Senior Sales Engineer, Corelight | Sep 24, 2026, 4:58:39 PM

Recent cyberattacks against U.S. water and wastewater systems have put some familiar operational technology (OT) security problems back in the headlines. Federal agencies have warned about malicious actors targeting internet-facing programmable logic controllers (PLCs), changing device configurations, and disrupting operations at utilities across multiple states.

What stood out was how little specialized OT knowledge was required to cause that disruption. In the reported activity, attackers could find internet-facing devices and gain access using default credentials that can often be found through a simple Google search. From there, they could change the device configuration. In some cases, changing an IP address was enough to interrupt communications with equipment that expected to find a device at a fixed address. Operators then had to respond to keep systems running. None of that requires a deep understanding of water treatment or industrial control systems (ICS), but it can still have a real operational consequence.

There are straightforward security recommendations that follow from that: Remove unnecessary internet exposure, change default credentials, segment OT networks, and harden devices. Federal guidance issued in response to the attacks makes many of those same recommendations. Anyone who has spent time around operational environments also knows that implementing them is far more complicated than listing them.

Security changes have operational consequences

Availability is one of the biggest constraints in OT. A water utility has to keep providing water, just as an electric utility has to keep power flowing and a manufacturing facility has to keep production moving. Taking equipment offline to patch, reconfigure, or replace it can interrupt the physical process the equipment supports. For a public utility, that can mean interrupting a service being provided to constituents.

The equipment itself adds another layer. Some of these environments rely on legacy or highly specialized devices that have been operating reliably for years. Some devices have their own routers or cellular connectivity, so operators can access them remotely rather than traveling to a site. Vendors may also have remote access for troubleshooting and maintenance. That connectivity exists for an operational reason, but it can also create direct internet exposure if it isn't managed carefully. Recent federal guidance specifically called attention to cellular modems installed by operators, vendors, or system integrators that may not be documented or captured in routine attack-surface scans.

This is a tradeoff OT security teams deal with constantly. Replacing an insecure device may require downtime. Removing a method of remote access can change how equipment is maintained. Smaller municipalities may also have limited security resources to work through these decisions. There are cases where hardening or segmentation can and should be done without much ambiguity, but there are others where improving security means figuring out how to build additional protections around equipment that can't easily be taken offline.

That makes understanding the environment itself especially important. One of the first things I would examine after incidents like these is what equipment is actually there, which devices have internet connectivity, and whether the segmentation you believe is in place matches what is happening operationally. If something is supposed to be isolated, I want to know what is actually communicating with it and how.

I've worked with organizations that told me a particular site was air-gapped. I would usually push on that a little and ask how people actually worked with the site. In one case, I asked someone on the physical security team how they retrieved security footage. Did they drive out there to get it? No, they just SSH'd in. There was a path from a corporate laptop with internet access into the supposedly air-gapped environment.

Situations like that aren't always obvious from an architecture diagram. Equipment gets added, vendors need access, and operational processes evolve. Asset inventories, segmentation, and internet exposure, therefore, need to be validated with network data to see how the environment is actually operating, particularly when that environment has been in place for a long time.

Understanding what normal looks like

OT environments have a characteristic that can work in a defender's favor: Much of what happens in them is repetitive. Processes run on schedules, devices communicate with the same systems, and operators become very familiar with how the equipment normally behaves. That creates a useful baseline when something changes.

A PLC moving into programmable mode is one example. There may be a legitimate reason for it during routine maintenance or a configuration change, but it shouldn't happen frequently. The same applies to a new command or an unexpected change in communications. The useful part is being able to see the change, establish when it happened, and put it in front of someone who knows whether that behavior makes sense for the process.

That collaboration matters in OT investigations. Engineers often know these systems extremely well but may not approach them from an attacker's perspective. Security practitioners bring that perspective, but they may not know everything that is possible or expected within a particular industrial process. Putting those two kinds of knowledge together helps distinguish a meaningful change from normal operations.

The evidence available to make that determination can also be different from what security teams are accustomed to in enterprise IT. Many operational devices can't run endpoint agents, which limits one of the most common sources of security telemetry. The devices still communicate across the network, though, so network activity can provide an independent source of evidence about what is happening between them.

The amount of context in that network data makes a difference during an investigation. Firewall logs and NetFlow may tell you which systems communicated, along with information such as ports and bytes. Once I'm investigating that connection, I want to know more. What protocol was being used? What happened during the session? How long did it last? How much data moved, and was a file transferred? Richer network evidence can help answer those questions. In an OT environment, understanding the protocol and commands involved can help connect network activity back to what a device was doing.

This gets harder when an attacker has legitimate credentials. A successful login to an exposed device may look very different from malware triggering a conventional alert. Understanding whether that connection is expected, what happened during it, and how the activity compares with the device's normal behavior gives an investigator more to work with.

What happened before operations were disrupted?

A lot of the conversation around attacks on critical infrastructure naturally turns to attribution. From a defender's perspective, attribution can be useful. Knowledge of how a particular adversary tends to operate gives threat hunters specific tactics and behaviors to look for in their own environments.

There are limits to how far that gets you. Adversaries change tactics, and every network gives them a different set of opportunities. The activity in these water-system incidents didn't require capabilities unique to a sophisticated nation-state actor. Understanding how access was obtained and what the attacker did with it gives organizations something concrete to examine.

There is still a piece of the recent water-system activity that I would like to understand better. We know when the operational impact became apparent, but the public information available doesn't establish how long attackers had access to individual devices beforehand. The configuration change that caused the disruption wasn't necessarily the beginning of the activity. The device first had to be discovered and accessed.

For the organizations affected, the answer depends in part on what data they were collecting at the time. For other critical infrastructure operators, the incidents offer a useful reason to examine that same issue now. If an unfamiliar system connected to an internet-facing operational device, what evidence would be available? Would it show only that a connection occurred, or enough about the interaction to determine what happened? How easily could that activity be compared with the way the device normally communicates?

That's the part of these incidents I keep coming back to: Whether the affected organizations had the data to see those connections before the attackers changed the IP addresses and created the outage. Once operations are disrupted, the problem is obvious. I'm more interested in what evidence was available before it got to that point.

Explore how Corelight provides network evidence for IT and OT environments.