All articles

Field notesAugust 27, 20269 min read

43 minutes without a link: what a node has to decide alone

We pull the link for forty-three minutes and count what a cloud-first security stack loses in the first ninety seconds. The list is longer than most architecture diagrams admit, and the hard part turned out to be the reconnect.

The short version

  • Seven controls at the site stopped making decisions during the outage. Not one of them reported that it had. Every dashboard stayed green.
  • Cloud EDR fell back to an offline policy that defaults to allow, and the identity provider locked operators out of the console they would have used to respond.
  • Five things have to be local: detection logic, operator authorization, the graph, pre-authorized response, and signed local evidence.
  • The reconnect is the harder engineering problem: priority drain, digest before payload, jittered backoff, and claims with provenance instead of an automatic merge.
43 minwith zero reach-back
0detections missed
0cloud calls required
4 sto drain the queue on reconnect

At 21:03:41 the SATCOM terminal drops. The outage is controlled and ends on a signal, but nothing in the security stack is told it is a drill. It lasts forty-three minutes. The timeline below is the denied-link rehearsal we run against power-site-03, the reference environment used throughout this blog, and it exists to answer one question: which of these is still making decisions?

The answer is the reason Kybernao Node exists, so treat the conclusion as motivated. The order in which things fail is not something we chose, though. It falls out of where each control keeps its brain.

Minute zero to ninety seconds

Very little announces itself. Almost nothing goes red. The degradation is quiet, staggered, and mostly invisible from the console the operator is looking at.

WAN DOWN 21:03:41DECIDINGNOT DECIDINGt+0+30s+90s+5m+15mcloud EDR verdictst+0sidentity providert+9sSIEM correlationt+31sthreat intelligencet+58snetwork scoringt+90slicense heartbeatt+5mticketingt+15m43 MINUTES · DETECTIONS MISSED: 0 · CLOUD CALLS REQUIRED: 0
Fig 1Nothing goes red. Each control keeps reporting healthy and stops deciding at its own moment.
ElapsedWhat changedWhat it looked like locally
t+0sCloud EDR verdict lookups failCached hashes still resolve. Unknown binaries fall to whatever the offline policy says, and that setting is worth reading before you need it rather than after.
t+9sIdentity provider unreachableExisting sessions live on. Token refresh fails. The console an operator would use to respond is the first thing to lock them out.
t+31sSIEM ingest buffersCorrelation rules spanning several sources stop firing, because half the sources stopped arriving. No rule reports that it is now blind.
t+58sThreat intelligence goes staleHarmless over 43 minutes. Serious over nine days, which is the actual planning case.
t+90sCloud-side network scoring returns nothingThe sensor still captures traffic. Nothing scores it. The dashboard shows a healthy sensor.
t+5mLicense heartbeat window opensVendor behaviour diverges sharply here. Read-only, fail-open and keep-enforcing are all defensible choices, and the one you bought is worth knowing.
t+15mTicketing unreachableThe tooling is fine. The process has nowhere to record a decision, so decisions stop being recorded.
The failure mode is rarely that a tool stops. It is that the tool keeps running, keeps reporting healthy, and quietly stops deciding.

That is what makes link loss a security event rather than an IT event. An adversary who can influence when your link drops can choose the ninety seconds in which most of your detection logic becomes a packet recorder.

What has to be local

The useful question is not whether a product "works offline". It is which five things it can still do alone. Our list, in the order that hurts to lose:

  • Detection logic and models. A detection that needs cloud scoring is not a detection at the edge. Models ship to the node and run there, including the expensive ones.
  • Authorization for operator actions. There has to be a local authority that can validate an operator and authorize a response without reach-back. In practice that means short-lived credentials pre-issued for the deployment, with a defined offline validity window that someone deliberately chose.
  • The graph. You cannot triage a path you cannot see. The terrain graph is resident on the node, not queried from a service.
  • Response authority. Containment rules are pre-authorized, with the conditions under which they may fire written down in advance. There is nobody to approve anything at 21:04.
  • Evidence. What happened has to be written, ordered, and signed locally, or you will not be able to reconstruct the outage afterwards, and the outage is exactly the window somebody will ask you about.
INSIDE THE BOUNDARY, OR IT DOES NOT DECIDEMISSION BOUNDARYdetection logicoperator authorizationterrain graphresponse authoritysigned local evidenceCLOUDenrichmentretentionNOTHING ON THE LEFT ASKS THE RIGHT FOR PERMISSION
Fig 2The cloud keeps its jobs. It just does not sit in the path of a decision.

Store-and-forward is a design problem, not a queue

"We buffer and send later" is where most designs stop. The details are where the data loss lives.

What to keep and what to drop

Bandwidth after a link returns is a fixed budget, so the drop policy has to be decided before the outage, not during the drain. We keep detections, evidence, and graph deltas. We drop raw capture. Graph state is queued as deltas rather than snapshots, because a snapshot is enormous and a delta compresses to almost nothing.

Bounded queues that tell the truth

Every queue is bounded, with priority classes, and overflow drops the lowest class first. The part that matters: the drop is itself recorded as an event, with counts by class. A silent drop turns your evidence record into a document that is confidently wrong about a window nobody can go back to.

Time, without a time source

Without NTP, clocks drift, and disconnected sites drift independently. Where a disciplined source is available we use it. Where it is not, everything carries two stamps: a monotonic local clock, and a sync-epoch marker so post-hoc correlation knows which segments are absolute and which are relative to one another. Choosing a single wall-clock stamp and hoping is how two nodes end up telling contradictory stories about the same minute.

Identifiers that survive a merge

Anything generated during the outage needs an identity that will not collide on reconnect. Node identity plus a monotonic counter; never an autoincrement from a shared sequence. This sounds like a database detail until two nodes federate and you discover you have two KYB-0041s.

The reconnect is harder than the outage

Forty-three minutes of backlog met a 9.6 kbps link. Three problems show up immediately, and only one of them is bandwidth.

Order. Drain FIFO and the operator receives the least important thing first, for several minutes, while the thing that mattered sits behind it. We drain by class: open findings, then response actions taken, then evidence digests, then bulk telemetry. Within a class, newest first, because a stale alert is worth less than a fresh one.

Shape. Digest before payload. The node sends a 200-byte summary and a content hash, and the far side asks for the payload only if it wants it. On a constrained link this is the difference between a useful picture in four seconds and a complete one in nine minutes.

QUEUED AT 21:47:10open findings3response actions taken1evidence digests12bulk telemetryrest9.6 kbps4sto everything an operator needed to seedigest 200 B first · payload on requestDRAIN ORDER · FINDINGS → ACTIONS TAKEN → EVIDENCE → BULK, NEWEST FIRST WITHIN A CLASS
Fig 3Bandwidth is the easy part. Order and shape decide what an operator sees in the first four seconds.

Everyone at once. When a shared uplink returns, every node behind it reconnects simultaneously. Jittered backoff and a per-node rate budget, or the first thing your restored link does is congest itself.

There is also a correctness problem underneath. After a partition, two nodes can hold state that genuinely disagrees. FedSpace does not merge them into a single truth; it exchanges claims with provenance attached, so the receiving side can see that node-a believed one thing at 21:20 and node-b another, and a human can adjudicate. Automatic merges hide exactly the disagreements worth looking at.

What the node actually did

During the forty-three minutes, with zero reach-back:

power-site-03/link.logNode
21:00:00  WAN up     federating selected state       sync 9.6 kbps
21:03:41  WAN down   SATCOM lost, store-and-forward  queue 0
21:04:02  Red        replayed KYB-004 against twin   no reach-back
21:04:19  Sentinel   observed 3/3 write vectors      detections local
21:04:33  Guardian   held containment on ot-dmz-fw-01 queue 3
21:47:10  WAN up     3 findings drained in 4s        sync 12.1 kbps

Detections missed: zero. Cloud calls required: zero. Drain time on reconnect: four seconds for everything an operator needed to see.

What we did lose is worth stating plainly. Two observations went without external threat-intelligence enrichment; the context arrived forty-three minutes late and changed neither verdict. One vendor reputation lookup never happened, and its answer would have been stale anyway. And a signature update sat in a queue on the other side of the link, which is fine at forty-three minutes and would not be fine at nine days. That is a real limit of local-first operation, not one we can engineer away.

Five rules we now design to

  1. Treat the link as a feature, not a foundation. Anything that cannot make its decision locally is a convenience, and should be labelled as one on the architecture diagram.
  2. Test by unplugging. Datasheets describe offline behaviour optimistically and almost never describe the default. Pull the cable in a controlled window and watch which dashboards keep claiming to be healthy.
  3. Pre-authorize the response. If the answer to "who approves containment at 21:04" is "the SOC," there is no answer at the edge.
  4. Log the degradation itself. "Threat intel stale for 43 minutes" and "EDR served cached verdicts for 108 unknown binaries" are security events. Most stacks record neither.
  5. Rehearse the reconnect. Plenty of systems survive the outage and fall over on the drain. That is the part nobody practises.

The measure of success is not the log above. It is that an operations team would have no reason to notice anything had happened. That is the whole goal: the mission does not pause because a satellite did.

Priya Natarajan Field Engineering

Runs deployment and denied-link rehearsal for Kybernao edge and expeditionary sites. Spends most of the year in places where the planning assumption is that the satellite will not be there.

Site design sessions open

Find out what your stack decides alone

Bring one site or one tactical mission. A Kybernao engineer will walk the controls with you and show which ones keep deciding when the link is gone.

Site design sessions open

Request a Kybernao demo

Tell us a little about your environment. A Kybernao engineer will follow up to arrange a focused walkthrough.

  • 01Focused architecture walkthrough
  • 02Mapped to your mission environment
  • 03Led by a Kybernao engineer
SECURE INTAKE→ENGINEERING
01Contact coordinatesRequired fields
02Mission profileFor a focused session

By submitting, you agree that Kybernao may contact you about this request.