
Jori VanAntwerp
For over two decades, Jori has enabled industrial and IT organizations to be successful in reducing risk, increasing compliance, and improving their overall security efforts. He has had the pleasure of working with companies such as Gravwell, Dragos, CrowdStrike, FireEye, McAfee, and is now CEO & Founder at EmberOT, a cybersecurity startup focused on making security a reality for critical infrastructure.
On the morning of Monday, July 27, the operating controls at the well and water treatment plant in Braham, Minnesota, went dark. Braham is a city of about 1,700 people. Public works crews had the plant running again in roughly two hours on manual operation.
Braham was one of more than 30 community water systems in the state hit in a coordinated cyberattack that weekend. On July 30, the FBI and EPA issued a joint public service announcement confirming that water and wastewater utilities in at least seven states had reported incidents, with actors remotely changing PLC IP addresses and passwords and operators losing monitoring and control.
CISA issued its own alert the same day, noting that across the wider campaign the activity had led to boil water notices and sustained manual operations, and that entities of all sizes were being targeted.
Nearly every piece written about it since lands in the same place: get your PLCs off the public internet. That’s correct, and I won’t argue with it.
The part I keep coming back to is what the operators did about it. What they had, whether formal or informal, was the foundation of an OT manual operations plan: people who knew the process, local controls that still worked, and enough procedure to keep water moving without supervisory control.
In Plymouth, communications dropped at two water towers and several lift stations, all of it cellular-connected equipment. The city pulled that gear off the network while crews kept the system running manually. South St. Paul switched to manual while its automated controls were restored. Maple Plain declared a local state of emergency to speed its response. State officials reported water quality was unaffected at every system hit, and no boil water advisories were issued anywhere in Minnesota.
Some of those utilities lost supervisory control. Plymouth gave it up on purpose, cutting off equipment it could no longer trust. Either way the water kept moving, because people who understood the process stepped in and worked it by hand.
In the incident write-ups, that shows up as a footnote. Manual operation gets logged as a degraded state, the sad middle of the timeline before automation comes back. It’s one of the strongest defenses available in an industrial environment, and OT is the only place you get it.
Why Manual Operation Works
Manual operation is the practice of running a physical process through local controls when supervisory systems are unavailable or can no longer be trusted. What makes it a defense comes down to what an attacker in an OT environment is usually after. The objective is a physical outcome: a pump that runs when it shouldn’t, a setpoint that drifts, a valve that moves, a plant that stops. Access to a controller is the road to that outcome.
Manual operation cuts the road.
When an operator takes local control of a pump station, the compromised path between the network and the physical process stops carrying anything that matters. Whatever the attacker still holds, and they may hold quite a lot, no longer reaches the thing they want to affect. The process is being driven by a person standing next to it, through a mechanism with no network interface to compromise.
That’s out of band in the most literal sense.
Hold that next to an IT environment. When a bank’s core systems go down, nobody clears transactions by hand. A hospital falls back to paper for a while, and even that’s partial and temporary. In OT the physical world is right there, waiting to be operated directly. The pump has a local control. The valve has a handwheel. The plant has an operator who knows what the flow should look like at 4 a.m.
Our industry spends a lot of time apologizing for OT. Legacy protocols with no authentication. Devices that can’t run an agent. Patch windows twice a year. All real. The physicality that makes so much of that hard also hands defenders something IT security has never had.
It isn’t available everywhere in OT, and I don’t want to oversell it. A high-speed packaging line or a continuous chemical process runs to tolerances no person can hold. Plenty of modern generation can’t be operated off the panel at all. The fallback is real wherever the process moves slower than people do, which covers a great deal of water, wastewater, distribution, and pipeline operations.
The Adversaries Worked This Out Before We Did
December 2016. Malware later named CrashOverride hit a transmission substation north of Kyiv and cut power to roughly a fifth of the city. It opened circuit breakers and held them open, so remote close commands went nowhere. Operators switched to manual and had power back in about 75 minutes. Locking them out of the remote path bought the attackers a little over an hour.
A year earlier, in December 2015, three Ukrainian distribution companies lost power to roughly 225,000 customers. Control center staff couldn’t regain remote control of more than 50 substations. Technicians drove out and closed the breakers by hand.
Then look at what those 2015 attackers built, documented in the SANS and E-ISAC analysis. KillDisk to wipe workstations. Scheduled UPS outages. Thousands of automated calls flooding the utility’s call center. And custom firmware for the serial-to-ethernet gateways at the substations, uploaded so that recovered operator workstations still couldn’t reach the field.
Every one of those targets the recovery rather than the outage. A well-resourced adversary spent months inside those networks learning how the utilities worked, then put real engineering into stopping them from coming back.
It wasn’t enough. Bricked gateways stopped remote commands and did nothing about a technician standing at the substation.
That’s the most durable field test we have of this idea, and it’s more than a decade old.
An OT Manual Operations Plan Has to Exist Before You Need It
An OT manual operations plan documents how operators will safely run a physical process through local controls when automated, remote, or supervisory systems are unavailable or untrusted. It is available on the day you need it only if the conditions already exist. Every utility that dropped to manual last month was drawing on decisions made years earlier, most of them by people who weren’t thinking about cybersecurity at all.
Those conditions double as the maintenance list.
Physical actuation still exists.
Somebody kept the local controls, the handwheels, the manual bypasses through the last modernization project. This is the one that quietly erodes. Every automation upgrade is a chance to delete a manual path because it looks like dead weight on a drawing.
People know the process itself.
What normal pressure feels like, which tank fills slowly in summer, what the chemical feed should be doing. That knowledge lives in operators and walks out the door when they retire.
The procedure is written down and current.
Manual operating procedures have a way of describing a plant that got reconfigured in 2019. If the last revision predates your last capital project, what you have is a document.
Somebody has practiced.
Your team’s first run at operating by hand shouldn’t happen at 2 a.m. with the console dark.
The chain of authority is clear.
Who declares manual mode. Who can authorize a physical change without the normal approvals. Hesitation on that question during an incident costs more than most technical controls save.
You know how long you can hold it.
Manual runs on people, and people have a duty cycle. Coverage that works for a shift may not work for a week. I’ve yet to see anyone publish a credible figure for how long a small utility can sustain it, which probably means each of us has to work out our own number before we need it.
None of that shows up as a security budget line today. We fund detection, we fund segmentation, we fund compliance evidence, and we treat the ability to run without any of it as an operational nicety. It belongs in the resilience column with a number next to it.
An OT Manual Operations Plan Needs a Return Path
The harder call comes later, when it’s time to hand control back.
At some point somebody decides the environment is trustworthy enough. That decision runs on evidence, and the evidence exists only if you captured what normal looked like beforehand.
Questions worth answering first:
- Do we know every path into the control environment, including the ones nobody documented?
- Have credentials and device configurations been restored from a known good state and verified?
- Do the current setpoints and project files match what we expect, or only what the screen says?
- Are devices talking to the same partners, on the same protocols, in the same patterns as before?
- If something is still in here, would we see it, or would we be trusting the same view that missed it the first time?
That last one is the whole game. Returning to automation means returning to trusting your instrumentation. If the only thing that changed is that the attacker went quiet, you’ve resumed rather than recovered.
CISA’s alert flags one item that belongs on this list: the exposed assets it’s seeing include cellular modems installed by operators, vendors, or system integrators that were never documented or picked up by routine attack surface scans. Mature programs included. Plymouth was pulling exactly that class of equipment off its network. You can’t scan for a modem somebody plugged in during a maintenance visit. You find it by watching what’s actually communicating inside your environment, which is the same visibility that makes the return decision defensible in the first place.
At minimum, an OT manual operations plan should answer:
- Which processes can be operated manually?
- Which local controls, valves, handwheels, or bypasses are required?
- Who can declare manual mode?
- Who is authorized to make physical changes?
- How long can the team sustain manual operations?
- What evidence is required before returning to automation?
The Small Operator Advantage
The utilities with the deepest manual capability are often the smallest ones. A system with three people, mechanical actuation throughout, and operators who’ve kept it running through storms and outages has a fallback that a fully automated regional operation may have engineered away. Every layer of automation that removes a human from the loop removes somebody who could have picked up the controls.
Automation is why these systems run as reliably as they do, and I’d build more of it. Just keep the manual path alive as you go.
The utilities that stayed up last month did so because people who understood their process could operate it without a network. That’s held in each of the events above, and the adversaries priced it in before our budgets did. Protect that capability. Write the procedure down, practice it, keep the local controls, and keep the knowledge in more than one head. You may never need it, but it’s a remarkable thing to have.
No noise. Just signal.
~Jori 🤘🔥
Become a Subscriber
EMBEROT WILL NEVER SELL, RENT, LOAN, OR DISTRIBUTE YOUR EMAIL ADDRESS TO ANY THIRD PARTY. THAT’S JUST PLAIN RUDE.
