The short version
- “Don’t touch OT” is the absence of a safety policy, and it leaves the route that matters, corporate identity into control, permanently untested.
- L0 Observe proves what the configuration permits. L1 Twin proves the design. Only L2 Live proves the deployment, because configuration is an intention and enforcement is a behaviour.
- Twins lie in three predictable ways: aged hardware, undocumented cabling, and drift between the exported config and the running device. A twin result is a strong claim about the design and a weak one about the plant.
- A live run is one authorized action, gated on process state, with abort conditions computed before the window opens and both owners’ signatures in the evidence record.
Ask most industrial operators whether their plant can be penetration tested and you will get one of two answers. No, or "in the November maintenance window." Both answers have the same consequence: the offensive picture of the environment is eleven months stale, and the route that actually matters, the one that crosses from corporate identity into control, never gets tested at all, because it spans the boundary where both teams' scopes stop.
"Don't touch OT" is not a safety policy. It is the absence of one. A safety policy states what may be done, by whom, under which process conditions, with which abort criteria, and who restores what afterwards. Writing that policy is most of the work. Kybernao Red implements it as three execution levels, and nothing runs outside them.
The three levels
| Level | What runs | Evidence produced | Authorization |
|---|---|---|---|
| L0 Observe |
Passive collection, configuration and identity analysis, graph reasoning. Nothing is sent toward a control device. | Reachability and trust claims derived from the environment's own configuration. | Standing, for the life of the deployment. |
| L1 Twin |
Full adversary emulation against a site digital twin built from the site's own configuration, captures, and controller logic. | A reproduced path with a replay set that can be run again on demand. | Security owner, per campaign. |
| L2 Live |
Non-destructive validation against the live environment. One authorized proof at a time, gated on process state, with pre-computed abort conditions. | Whether the real boundary permits or denies the action, on the real equipment. | Two-party: security owner and process owner, inside a named window. |
A campaign walks up the levels and stops at the first one that answers the question. Most findings never need L2. The ones that do are usually the ones a defence depends on.
L0: what you can prove without sending a packet
A surprising amount is provable statically. Reachability derived from exported firewall rule sets, routing and switch configuration, cloud role assignments, and the identity graph is a genuine proof about what is permitted. On the seven-hop path we published earlier this year, L0 alone produced hops one through five.
What L0 cannot tell you is whether the deny you believe in actually enforces. Configuration is an intention; enforcement is a behaviour. Devices diverge from their documented rule sets, vendors interpret their own syntax creatively, and shadowed rules do not announce themselves. That gap is the entire reason levels 1 and 2 exist.
L1: twin fidelity, and where twins lie
A twin is only worth what it reproduces. Ours are built from the site's own artefacts, and a twin is not accepted for validation until it carries all five of these:
- The protocol stack as deployed, including the vendor's deviations from the published specification, which are the interesting part.
- Controller logic: the program and tag map, imported from the project file, not inferred from traffic.
- The timing envelope: scan cycles, watchdog timeouts, session lifetimes. A technique that works at 50 ms and fails at 2 ms is a different result, not the same result faster.
- Boundary device rule sets exported from the running devices, never transcribed from documentation.
- The safety instrumented system's trip points, modelled as boundaries the twin refuses to cross, so a test cannot quietly "succeed" by driving the model somewhere the real plant would have tripped.
And then the honest part. Twins lie in three predictable ways.
They lie about aged hardware: a fifteen-year-old controller under thermal and cycle load behaves differently from a faithful simulation of its manual. They lie about what nobody documented: the twin has no forgotten serial line, no contractor's laptop left on the engineering VLAN, no half-decommissioned wireless bridge. And they lie about drift: the twin inherits the configuration you exported, so if the running device has diverged from it, your twin models your paperwork rather than your plant.
A twin result is a strong claim about the design and a weak claim about the deployment. That distinction is the whole reason we do not stop at the twin.
L2: non-destructive proof, and abort conditions
Live validation is narrow on purpose. The rules below are not guidance; they are enforced by the execution engine and recorded in the authorization record.
- One proof per authorization. The exact actions, meaning packets, sessions and commands, are written into the authorization before it is signed. Anything not in the record does not execute.
- Prefer the negative proof. Attempting an action and showing it is denied is safe, repeatable, and usually more valuable than performing it and showing that it worked.
- Read before write. Where a write path must be proven and a spare register block with no process effect exists, that is the target. Where none exists, the finding stops at L1. It does not get promoted because somebody is curious.
- Single attempt. No brute force, no fuzzing, no sweeping scans of control segments. Broadcast discovery has knocked over controllers that no exploit could touch.
- Process-state gating. The window is defined by plant state, not only by the clock. If the process leaves the declared state, say a unit comes online or a transfer starts, the run aborts even if there are twenty minutes left on the authorization.
- Abort conditions computed up front. Latency on the control segment above a threshold, controller CPU or scan time above a threshold, a watchdog counter moving, any safety system annunciation, or an operator's hand on the stop. Any one of them aborts and triggers the restore plan.
Everything executes on the node and is logged there against its authorization identifier, so the plant's own record, not a vendor's cloud, shows exactly who allowed what, when, and what happened.
The result worth having
The most valuable live outcome is almost always a denial. "The unsigned write is refused at the boundary" is the claim your entire defence rests on, and at most sites it is the claim nobody has ever tested.
Who says go
L2 requires two parties, and the record they sign is a document rather than a checkbox. It names the finding, the route, the exact action, the window, the process state required to begin, the abort conditions, the restore plan, and both signatories. In outline it looks like this.
authorization KYB-004-L2-014
finding KYB-004 unsigned modbus write -> PLC-07
action single fn=16 write, spare block 40100-40101
window 02:00-02:40, one named night
state unit 3 offline, aux cooling on bypass
abort scan_time > 12ms | sis_annunciation | operator_stop
restore revert 40100-40101, confirm with engineering
signed security_owner + process_owner, both requiredThat record is part of the evidence package, not the ticketing system. When a regulator asks what was done to that controller on the nineteenth, the answer is one query against the node, signed, with the log of what actually executed attached to it.
Closing the loop
Levels apply to retests too, and retests are usually cheaper than the original discovery. Discovery often has to run at L1 because proving a write path is intrusive. The retest is a denial: the write is attempted and refused, which is safe enough to run live, on a schedule, without a maintenance window.
That inversion is the point of the whole system. Finding the path once is a project. Knowing on any given Tuesday whether it is still closed is an operating capability, and it is the one that actually reduces risk, because configurations drift and people edit rules for perfectly good reasons.
What we will not do
Shorter list, and it does not move:
- No fuzzing of control protocols on live segments. Twin only, always.
- No denial-of-service testing against a production process. There is no version of that with an acceptable blast radius.
- No credential spraying against an authentication service that can lock out operator accounts. Locking out the control room is an incident, not a finding.
- No persistence on control devices. We do not leave anything on a controller, ever.
- No testing during start-up, shutdown, black start, planned islanding, or any operation the site has declared critical.
The last item is not a technical limit. It is the sort of line that gets written by someone who has stood in a control room during a start-up, which is why the policy is co-signed by a process engineer rather than drafted by security alone.
None of this makes offensive testing safe in the abstract. It makes it bounded, authorized, and repeatable, and it means the answer to "can you test our plant" stops being November.