Read the Gauge, Not the Screen
Twelve states, no patch coming, and the question the water sector should be asking instead
Two weekends ago, more than thirty Minnesota communities had the control systems on their water and wastewater plants attacked at the same time.
In Braham — a town of about 1,700 people — the attack knocked out the computerized controls and shut down the well and the treatment plant. Public works crews walked in and had it running again in about two hours. In Plymouth, the city’s IT staff pulled the cellular-connected equipment at two water towers and a string of lift stations off the network, and the crews ran the system by hand. South St. Paul kept water flowing. Maple Plain declared a local emergency and kept water flowing.
Nobody in Minnesota got a boil-water notice. Water quality was never affected there.
Then it kept going. Michigan. South Dakota. New Jersey. Georgia. By August 4th, utilities in at least twelve states were working on incidents. In Clayton County, Georgia — a system serving about 300,000 people outside Atlanta — the attack dropped water pressure far enough that the authority issued a precautionary boil-water advisory. It was lifted after testing. And in its July 30th announcement, the FBI reported that operational effects across the campaign have included loss of pressure and flooding and noted that pressure loss can allow untreated groundwater to seep into pipes.
That last sentence is the federal government describing a path from a keyboard to your tap.
The crews have done well. Nothing has been contaminated. But the industry has been leaning on an assumption for thirty years — when the automation goes down, you go manual, and manual always works — and I want to argue that the floor is not where we think it is. Then I want to make a larger argument about how we decide where floors are in the first place.
The part that didn’t make the news
Four days before Minnesota, CISA updated a federal advisory covering this campaign. Down in the detection guidance is an observation the FBI made at one victim, and it is the most important thing published about any of this.
The attacker got into the controller and did two things.
First, he modified the reusable code modules inside the controller program to disable the safety shutdowns and the alarms. Everything downstream of those modules looked normal. An engineer opening that program and reviewing it the way programs get reviewed would have seen exactly what he expected to see.
Second, he manipulated the data showing up on the HMI and SCADA screens.
Read those two together and sit with it. The protection that stops an unsafe excursion: removed. The indication that would tell a human being the excursion was happening: falsified. Equipment can run in an unsafe condition, and nothing in the control room says a word.
Now go back to “we’ll just go manual.”
Manual isn’t magic
Manual operation is not a magic state where the plant becomes safe. It is a human being making decisions about a physical process based on what he knows about that process.
Where does he get what he knows? Off the screen. Off the trend. Off the same instrument path that just got corrupted.
Going manual against falsified indication does not restore control. It hands the wheel to a man who has been given a wrong map and no reason to doubt it.
This is worth being precise about, because the two failures are not cousins. Losing your indication is a survivable failure. An operator staring at a dead screen knows he is blind, and behaves like a man who is blind — he slows down, he walks the plant, he calls someone. Every manual backup procedure ever written assumes that condition. And that is what most of these utilities faced: the FBI describes the common effect as loss of view, operators locked out of equipment they could no longer see or command.
Falsified indication is the opposite condition. He is not blind. He is confident, and he is wrong, and there is nothing anywhere in his environment that disagrees with him. That is the rarer case in this campaign. It is also the one that matters for design, because a plant built to survive the harder case survives the easier one and the reverse is not true.
Here are the findings, and they are not comfortable. Automatic protection and manual fallback are supposed to be two independent layers of defense. If an attacker can corrupt what the operator sees, they are not two layers. They are one layer wearing two hats — and one modification takes both.
There is no patch, and this time there’s no patch coming
The FBI named the equipment: Allen-Bradley MicroLogix 1100 and 1400 controllers, reachable from the open internet. The attackers logged in, changed the IP addresses, and set passwords. That’s it. No exploit chain. They locked the owners out of their own equipment by configuring it.
And then look at what the FBI put in the mitigations section — a full block on end-of-life hardware. Maintain a rolling twelve-month EOL forecast. Track EOL systems by product, owner, and location. Replace or isolate them, with firm decommission dates.
They put that in there for a reason. When a product line ends, the vendor stops selling it, stops supporting it, and stops issuing security patches. Forever. There is no update coming, not because the flaw is too deep to fix, but because nobody is going to fix anything on that box again.
Sit with that, especially if your security program is built the way most water utility security programs are built. Scan, prioritize, patch, report compliance, repeat.
That program has no move here. There is nothing to install and there never will be. Your options are to change the architecture around the box or to buy a new box. Nothing else on the list is real.
Which raises the question this whole essay is actually about. If the fix isn’t a fix, what do you do instead?
A different question to start from
Almost every security framework our sector uses starts from the adversary. Who might come after us, how capable are they, what could they do, which controls address that? IEC 62443 grades systems by the capability of the attacker they can resist. Risk assessments rank findings by likelihood and impact. Vulnerability programs work a list of known flaws.
That is a sound way to think, and I want to be clear that I am not throwing rocks at it. AWIA moved this sector further in five years than the previous twenty. J100 is a serious methodology. IEC 62443 is a good framework for what it governs. They were all built to answer the question “how do we manage risk from an adversary?” and they answer it.
A design basis starts somewhere else entirely. It asks:
What must never happen here — and what has to remain true for that to hold, no matter who is on the other end?
Notice what that does. On the adversary axis, you are maintaining a list that somebody else writes. You are always one move behind, by construction, because the other guy picks the next entry. On the consequence axis, the adversary does not get a vote. You define the envelope of conditions you must survive. You demonstrate — in advance, on paper, where colleagues can argue with you — that you survive every point in it. Then you go find out whether you were right.
This is not a new idea, and I did not invent it. It is how commercial nuclear power has been licensed since the 1960s, and it is the reason that industry can tell you what happens in an accident it has never had. Three rules make it work, and all three land squarely on what happened these past two weeks.
The demonstration cannot rest on anything an adversary can take away from you. This is where Nuclear’s obsession with passive features comes from — gravity, natural circulation, thermal mass, the geometry of the thing. Features that keep working when the power is out, the operators are gone, and the network is lying. That single constraint is why the safety case survives conditions nobody enumerated in advance. Apply it here, and it cuts immediately: your safety case cannot rest on the SCADA screen telling you the truth, because a screen is exactly the kind of thing that can be taken away from you.
Layers only count if they fail independently. Everybody says defense in depth. Most people mean “we have several controls.” A design basis makes you prove the layers don’t share a common failure — that no single thing takes two of them at once, which is precisely the finding above. Automatic protection and manual fallback both consume the controller’s version of reality. On the org chart, two layers. In fact, one.
And it runs bigger than one plant. The FBI noted that across several victims, similar network setups supplied by third parties may let the attackers multiply their successes. Think about what that means. Thirty systems in one weekend may not have been thirty break-ins. It may have been one architecture, understood once, and used thirty times. If your integrator built your neighbor’s plant the same way he built yours, you and your neighbor are not two independent utilities. You are one configuration with two addresses.
It gets written down once, and inherited. The analysis is done for a class of facility, reviewed once, and referenced by everyone who operates that class. Nuclear does this with topical reports, and it is the reason a small operator doesn’t have to fund a first-principles safety case. A town of 1,700 should not be expected to invent its own. It should be able to point at one and say which parts apply to it.
What that buys you, right now
Three payoffs, and the last two weeks are the proof of all three.
It survives the un-patchable. A program organized around remediation dies the moment there is nothing left to install. A design basis never depended on the patch, because it never asked: “is this flaw fixed?” It asked, “can an unauthorized party change this controller’s behavior?” That question still has an answer when the vendor has stopped answering the phone.
It reaches things nobody has enumerated yet. A requirement that says credentials for control system access must not be replicable in software catches a whole family of problems without anyone involved knowing which specific box will be attacked. A control list keyed to known vulnerabilities cannot do that, structurally. It can only ever contain what somebody already found — and the FBI’s advisory on this campaign doesn’t cite a CVE at all, because there isn’t one to cite.
It tells you when you’re done. Compliance tells you what you did. A design basis tells you whether it was enough, because adequacy gets measured against the consequence you refused to accept, not against the checklist you completed. That is the whole distance between “we passed our audit” and “we are actually covered.”
Unknown risk is the worst position available in this business, because you cannot manage, transfer, or insure what you have never characterized. The design basis forces the characterization before the event instead of after it. That is the entire value proposition, and it is not a compliance argument — it is a not-being-surprised argument.
About the small systems
Braham serves about 1,700 people. AWIA — the federal standard governing water system risk assessment and emergency response planning — applies to systems serving 3,300 and up. Braham is below the line. So are most of the roughly 150,000 water systems in this country.
I’ll correct something I wrote about this a week ago. I said the attacker was going after the small systems on purpose, picking the segment our standards excused. The fuller picture doesn’t support that. Clayton County serves 300,000 and got hit in the same stretch as Braham. What the attacker is selecting for is a particular kind of controller, hanging off the internet, on a particular kind of connection — and that cuts across system size, not along it.
But something else in the record cuts the other way, and it’s worse than what I originally claimed.
The one organization known to have caught the modified controller logic found it by noticing that the ladder logic didn’t match across several of its own sites. It compared its plants to each other. That is a fine detection method if you run eight plants. If you run one plant in a town of two thousand, you have nothing to compare against. The small system isn’t necessarily the preferred target. It’s the one that can’t see, can’t resist, and can’t recover — and it’s outside the scope of the standard written to protect the sector.
The answer isn’t a smaller checklist. It’s a design basis somebody else already wrote that a system serving two thousand people can adopt in an afternoon.
Six things, no budget required
So here is a design basis in miniature. Each item is an answer to “what must hold no matter who is on the other end.” None of it needs a security consultant, a product, or a line item — and every one of them now appears in federal guidance too.
1. No controller reachable from the internet. Not “behind a password.” Not reachable. And check the field sites — the tower, the lift station, the remote well on a cellular modem. Those are the ones nobody remembers exist, because they never touch the office network and never show up in an IT inventory. The FBI calls out cellular field connectivity specifically.
2. Physical mode switch set to block remote program changes. It’s a hardware switch on the controller, in the RUN position. Cost is a site visit. It is the difference between a controller a stranger can reprogram from another continent and one he cannot. One caution the FBI adds: check the program before you flip it to RUN, because flipping it locks in whatever is loaded.
3. An offline copy of your controller program that you know is good. Stored where the network cannot reach it. If you run one plant, this isn’t a nice-to-have — it’s your only way to answer “is the program running my plant the program I wrote?” And if you’re restoring from backup after an incident, check the backup first. A backup made after the attacker got in is not a backup.
4. A manual operating procedure somebody has actually run. Not a binder on a shelf. Practiced, on a schedule, by the people who would have to do it at two in the morning on a Sunday. The FBI listed the ability to switch to manual as one of the things that determined how badly each victim got hurt. Braham’s crew was not reading instructions for the first time.
5. Local indication you can read with your own eyes, checked against the screen on a schedule. A gauge. A residual analyzer with a face on it. A sight glass. Then the habit that makes it worth having: walk it down, compare it to what the SCADA screen says, write down the difference. And if the screen and the gauge disagree — believe the gauge, and start asking questions.
6. Know what’s end-of-life, and have a date on it. Walk your plant and write down every controller, every modem, every HMI, with its model and its support status. Anything the vendor has stopped supporting is a box that will never be fixed again. It doesn’t have to be replaced tomorrow. It has to be on a list, with a date, and isolated until that date arrives.
Number five is the one almost nobody has, and it’s the cheapest on the list. It’s the passive feature. It’s the thing the attacker cannot reach across the network and change.
The same argument is coming for AI
I’ll close with the reason this matters beyond water.
The same design basis reasoning is being worked right now against AI in control rooms, and the parallel is uncomfortably exact. The comforting phrase in that conversation is “there’s a human in the loop.” It gets said the way we say “we’ll just go manual” — as though naming the fallback establishes that the fallback works.
It’s the same floor, and it fails the same way. A human supervising a system whose outputs he cannot independently verify is not a safeguard. He is a person being told what to think by the thing he is supposed to be checking. The design basis question fixes it in one move, and it is the same question as before: what does that human have that does not come from the system he is supervising?
If the answer is “nothing,” you don’t have a human in the loop. You have a human in the display.
Twelve states, and the water is still clean. The crews earned that. I would much rather this sector learn the lesson from the version that went well than wait for the version that doesn’t.
For what it’s worth, I’ve spent the spring writing a water sector design basis on exactly this premise — that credentials for control system access must not be replicable in software, and that cyber protection has to be architectural rather than incident-specific because the threat does not resolve on anyone’s timeline. I’m not claiming a crystal ball, and I’d rather not have been right this way. I’m claiming the posture is the correct one. This month the vendor guidance and the federal advice both came out architectural, because nothing else was on the table.
Read the gauge.
The FBI and EPA public service announcement of July 30th, updated August 6th, is worth reading in full. It is short; it names the equipment, and the mitigations section is the best free security assessment a small utility is going to get.


A great example of how the nuclear sector has solved problems that other sectors have not faced, until now. Great work Sir Tim!