The short version
- Seven controls at the site stopped making decisions during the outage. Not one of them reported that it had. Every dashboard stayed green.
- Cloud EDR fell back to an offline policy that defaults to allow, and the identity provider locked operators out of the console they would have used to respond.
- Five things have to be local: detection logic, operator authorization, the graph, pre-authorized response, and signed local evidence.
- The reconnect is the harder engineering problem: priority drain, digest before payload, jittered backoff, and claims with provenance instead of an automatic merge.
At 21:03:41 the SATCOM terminal drops. The outage is controlled and ends on a signal, but nothing in the security stack is told it is a drill. It lasts forty-three minutes. The timeline below is the denied-link rehearsal we run against power-site-03, the reference environment used throughout this blog, and it exists to answer one question: which of these is still making decisions?
The answer is the reason Kybernao Node exists, so treat the conclusion as motivated. The order in which things fail is not something we chose, though. It falls out of where each control keeps its brain.
Minute zero to ninety seconds
Very little announces itself. Almost nothing goes red. The degradation is quiet, staggered, and mostly invisible from the console the operator is looking at.
| Elapsed | What changed | What it looked like locally |
|---|---|---|
| t+0s | Cloud EDR verdict lookups fail | Cached hashes still resolve. Unknown binaries fall to whatever the offline policy says, and that setting is worth reading before you need it rather than after. |
| t+9s | Identity provider unreachable | Existing sessions live on. Token refresh fails. The console an operator would use to respond is the first thing to lock them out. |
| t+31s | SIEM ingest buffers | Correlation rules spanning several sources stop firing, because half the sources stopped arriving. No rule reports that it is now blind. |
| t+58s | Threat intelligence goes stale | Harmless over 43 minutes. Serious over nine days, which is the actual planning case. |
| t+90s | Cloud-side network scoring returns nothing | The sensor still captures traffic. Nothing scores it. The dashboard shows a healthy sensor. |
| t+5m | License heartbeat window opens | Vendor behaviour diverges sharply here. Read-only, fail-open and keep-enforcing are all defensible choices, and the one you bought is worth knowing. |
| t+15m | Ticketing unreachable | The tooling is fine. The process has nowhere to record a decision, so decisions stop being recorded. |
The failure mode is rarely that a tool stops. It is that the tool keeps running, keeps reporting healthy, and quietly stops deciding.
That is what makes link loss a security event rather than an IT event. An adversary who can influence when your link drops can choose the ninety seconds in which most of your detection logic becomes a packet recorder.
What has to be local
The useful question is not whether a product "works offline". It is which five things it can still do alone. Our list, in the order that hurts to lose:
- Detection logic and models. A detection that needs cloud scoring is not a detection at the edge. Models ship to the node and run there, including the expensive ones.
- Authorization for operator actions. There has to be a local authority that can validate an operator and authorize a response without reach-back. In practice that means short-lived credentials pre-issued for the deployment, with a defined offline validity window that someone deliberately chose.
- The graph. You cannot triage a path you cannot see. The terrain graph is resident on the node, not queried from a service.
- Response authority. Containment rules are pre-authorized, with the conditions under which they may fire written down in advance. There is nobody to approve anything at 21:04.
- Evidence. What happened has to be written, ordered, and signed locally, or you will not be able to reconstruct the outage afterwards, and the outage is exactly the window somebody will ask you about.
Store-and-forward is a design problem, not a queue
"We buffer and send later" is where most designs stop. The details are where the data loss lives.
What to keep and what to drop
Bandwidth after a link returns is a fixed budget, so the drop policy has to be decided before the outage, not during the drain. We keep detections, evidence, and graph deltas. We drop raw capture. Graph state is queued as deltas rather than snapshots, because a snapshot is enormous and a delta compresses to almost nothing.
Bounded queues that tell the truth
Every queue is bounded, with priority classes, and overflow drops the lowest class first. The part that matters: the drop is itself recorded as an event, with counts by class. A silent drop turns your evidence record into a document that is confidently wrong about a window nobody can go back to.
Time, without a time source
Without NTP, clocks drift, and disconnected sites drift independently. Where a disciplined source is available we use it. Where it is not, everything carries two stamps: a monotonic local clock, and a sync-epoch marker so post-hoc correlation knows which segments are absolute and which are relative to one another. Choosing a single wall-clock stamp and hoping is how two nodes end up telling contradictory stories about the same minute.
Identifiers that survive a merge
Anything generated during the outage needs an identity that will not collide on reconnect. Node identity plus a monotonic counter; never an autoincrement from a shared sequence. This sounds like a database detail until two nodes federate and you discover you have two KYB-0041s.
The reconnect is harder than the outage
Forty-three minutes of backlog met a 9.6 kbps link. Three problems show up immediately, and only one of them is bandwidth.
Order. Drain FIFO and the operator receives the least important thing first, for several minutes, while the thing that mattered sits behind it. We drain by class: open findings, then response actions taken, then evidence digests, then bulk telemetry. Within a class, newest first, because a stale alert is worth less than a fresh one.
Shape. Digest before payload. The node sends a 200-byte summary and a content hash, and the far side asks for the payload only if it wants it. On a constrained link this is the difference between a useful picture in four seconds and a complete one in nine minutes.
Everyone at once. When a shared uplink returns, every node behind it reconnects simultaneously. Jittered backoff and a per-node rate budget, or the first thing your restored link does is congest itself.
There is also a correctness problem underneath. After a partition, two nodes can hold state that genuinely disagrees. FedSpace does not merge them into a single truth; it exchanges claims with provenance attached, so the receiving side can see that node-a believed one thing at 21:20 and node-b another, and a human can adjudicate. Automatic merges hide exactly the disagreements worth looking at.
What the node actually did
During the forty-three minutes, with zero reach-back:
21:00:00 WAN up federating selected state sync 9.6 kbps
21:03:41 WAN down SATCOM lost, store-and-forward queue 0
21:04:02 Red replayed KYB-004 against twin no reach-back
21:04:19 Sentinel observed 3/3 write vectors detections local
21:04:33 Guardian held containment on ot-dmz-fw-01 queue 3
21:47:10 WAN up 3 findings drained in 4s sync 12.1 kbpsDetections missed: zero. Cloud calls required: zero. Drain time on reconnect: four seconds for everything an operator needed to see.
What we did lose is worth stating plainly. Two observations went without external threat-intelligence enrichment; the context arrived forty-three minutes late and changed neither verdict. One vendor reputation lookup never happened, and its answer would have been stale anyway. And a signature update sat in a queue on the other side of the link, which is fine at forty-three minutes and would not be fine at nine days. That is a real limit of local-first operation, not one we can engineer away.
Five rules we now design to
- Treat the link as a feature, not a foundation. Anything that cannot make its decision locally is a convenience, and should be labelled as one on the architecture diagram.
- Test by unplugging. Datasheets describe offline behaviour optimistically and almost never describe the default. Pull the cable in a controlled window and watch which dashboards keep claiming to be healthy.
- Pre-authorize the response. If the answer to "who approves containment at 21:04" is "the SOC," there is no answer at the edge.
- Log the degradation itself. "Threat intel stale for 43 minutes" and "EDR served cached verdicts for 108 unknown binaries" are security events. Most stacks record neither.
- Rehearse the reconnect. Plenty of systems survive the outage and fall over on the drain. That is the part nobody practises.
The measure of success is not the log above. It is that an operations team would have no reason to notice anything had happened. That is the whole goal: the mission does not pause because a satellite did.