This is the fourth post in our MTSA cybersecurity series. In the first post, we introduced the crawl, walk, run approach we’re using to work through the regulation. In the second post, we covered asset inventory, and in the third, network segmentation. This post picks up with resilience.
You have your asset inventory. You’ve segmented your network. Now it’s time to build resilience, the capability that determines how well you survive what you can’t prevent.
Resilience used to be one of those words that meant everything, and therefore nothing at all; MTSA fixed that. The goal is to ensure that U.S.-flagged vessels, facilities, and OCS facilities can recover quickly from cyber incidents with minimal disruption to operations. The expectation is straightforward: cyber incidents should be prevented, detected, responded to, and recovered from, with as little operational impact as possible.
The regulation breaks resilience into four requirements: reporting, response, validation, and recovery. Four steps. Deceptively simple on paper. In the first post, we introduced crawl, walk, run as the way we work through every MTSA requirement in this series. Here’s what that progression looks like applied to resilience, across all four requirements.
At this stage, you’re not meeting the regulation yet; you’re building what it stands on. The goal is basic readiness, nothing more.
Response: Document incident response roles, even if informally. Identify critical OT and IT assets and their role in vessel or facility operations, and define minimum operational states: what must stay running versus what can fail.
Validation: Ensure basic network segmentation exists between IT and OT environments, and perform periodic backup validation, even if that validation is infrequent or manual.
Recovery: Establish basic backup procedures, manual or automated, for key systems. Develop manual fallback procedures for critical operations, including paper-based or local control alternatives for key functions, and make sure operators know what to do if primary systems go down.
At this stage, the deliverables are modest but real: an asset list with criticality designations, a basic incident response contact sheet, a documented minimum operating baseline, and manual fallback SOPs.
This is where the regulation’s actual requirements live. MTSA expects a formal, tested program, not a folder of good intentions.
Reporting: If you have a cyber incident, report it. Keep your response above board, see something, say something. Even if you are not currently bound by existing Coast Guard reporting rules, you are still required to immediately report cyber incidents that have met the reporting threshold. That threshold is any event that causes or could lead to a significant impact on operations, safety, or data confidentiality. The goal of cyber incident reporting is to establish a rapid reporting expectation that provides federal authorities with visibility into emerging threats. Think of it less as paperwork and more as an early warning system for maritime environments. The incident you have today could be someone else’s way of prevention tomorrow. The Coast Guard cannot assist or connect dots it doesn’t see.
Response: The regulation now specifies that a Cyber Incident Response Plan (CIRP) must be developed, implemented, maintained, and exercised as part of the overall Cybersecurity Plan. The CIRP is much more than a piece of paper; it is a living document within the Cybersecurity Plan. We have seen CIRPs created in 2018, never updated, and listing people who no longer work at the company. That document isn’t a plan; it is a liability. The CIRP must be actionable, updated regularly, and tested. The roles of key stakeholders, including the Cybersecurity Officer (CySO), must be defined to establish a clear communication path and the actions to be taken in the event of a cyber incident. Essentially, an IR playbook built for maritime operations, not a generic IT template. A CIRP reinforced with regular exercises such as tabletop scenarios, simulations, or live drills directly reduces response times when an actual incident occurs. Organizations that test their plans respond faster, with less confusion, and with fewer operational disruptions. That response capability rests on the underlying architecture: redundancy for critical systems across network paths, control systems, and communications, plus OT monitoring and detection that surfaces disruption signals early enough to act on. Don’t wait for a digital storm to find your sea legs – chart your course with this guide to OT Incident Response Planning.
Validation: The Cybersecurity Plan must be validated for effectiveness through annual testing, including exercises, incident response case reviews, and updates to controls based on lessons learned. The Cybersecurity Plan is not a document you can create and toss aside; it must be relevant, actionable, and tested. Seriously, we have walked into facilities where the Cybersecurity Plan was printed, signed, put into a pretty binder, and placed on a shelf, gathering dust alongside the operator manuals from 1982. That is not compliance, that is decoration. The tide doesn’t stand still, and neither can your Cybersecurity Plan. Incorporating new threats and lessons learned keeps your defenses current and reduces the risk of repeat incidents. The lessons learned are not siloed to your own; they also include those from third-party and supply chain incidents. If a critical vendor suffers a breach that impacts a system they manage for you, who makes the call to isolate that system? Whose responsibility is it to call the Coast Guard if the reporting threshold has been met? Supply chain incidents are not just theory in maritime environments. The Colonial Pipeline ransomware attack impacted fuel shipping and port logistics, and the Mediterranean Shipping Company ransomware disruptions caused impacts to terminal operations, required vessel diversions, and disrupted transmission of vessel bayplans, load lists, and customs information between shipping partners.
Recovery: Critical IT and OT-identified assets must be backed up, with the backups protected and tested regularly. To secure backups, they should be kept offline or protected using Write Once, Read Many (WORM) technology, which prevents tampering or compromise by making them tamper-resistant and immune to data modification or deletion. In addition, the backups should be validated and tested regularly. It is important to know that you can use and quickly restore data from your backups before the time comes when you must use them. Backups are your last line of defense; they are your fallback. Just like an actual life raft, you do not want the first time you test it to be when the ship is going down. An untested backup is just a file that gives you warm fuzzies. At this stage, RTOs and RPOs are defined and documented for critical systems, failover processes are validated in controlled test scenarios, backup restoration is tested against those defined RTOs, and exercise outcomes get reviewed and folded back into updated plans.
At this stage, the deliverables mature: documented RTOs/RPOs, a tested backup strategy, incident response playbooks, tabletop exercise records, and an OT monitoring baseline.
This is where resilience stops being a compliance exercise and starts being how the organization actually operates.
Response: Maintain a segmented, zero-trust-informed architecture across IT and OT, and integrate resilience into supply chain and vendor risk management, since the same third-party questions raised in validation need standing answers, not ad hoc ones during an incident.
Validation: Continuously test resilience through red team, purple team, and adversarial simulations. Use real-time monitoring and automated response for containment, and perform full-scale recovery exercises, not just tabletop ones.
Recovery: Implement automated failover and high availability for critical systems, and align resilience metrics to mission impact: port uptime, vessel safety, cargo flow. Automated detection-to-containment workflows with minimal human latency, with recovery outcomes feeding a continuous improvement cycle rather than a filed report.
At this stage, the deliverables reflect a program that runs itself: automated failover documentation, red and purple team reports, mission-impact resilience metrics, and full-scale exercise after-action reports.
Resilience only works if you’ve already done the first two steps. An asset inventory tells you what to protect. Segmentation tells you where the boundaries are. Resilience is what happens when both are tested against a real incident, and it only holds up if the next requirement is also in place: keeping all of it maintained. That’s routine system maintenance, the subject of the next post in this series.
· A Resilience Maturity Framework mapping OT/IT resilience across three progressive stages, aligned to federal cybersecurity requirements.
Need help building your OT incident response plan?