Kyber Cypher

Learn

Field Log 029

How does the watchdog know it is not the one that is broken?

Field Log // 029 Status Open Difficulty Free Cost Nothing
The Story

I put a watchdog on a phone. It sits in my pocket, checks that my machines are answering, and tells me when they are not. A monitor that lives somewhere else is a genuinely good idea, and it immediately taught me two things I did not want to know.

The first was a timing problem. The watchdog has a deadman timer: it is supposed to check in on a fixed rhythm, and a missed check in is itself the alarm. For days the rhythm was wrong. Not broken, wrong. Checks that should have been minutes apart were arriving tens of minutes late, in bursts, as though the phone had been asleep and then remembered.

It had. A modern phone aggressively suspends background work to save battery. It does not tell your timer it has been postponed; it simply postpones it. So the deadman timer, whose entire job was to be reliable about time, was running on a device that treats time as negotiable. The fix was to have the check take a wake lock for the few seconds it needs, so the operating system keeps the processor awake until the work is genuinely done.

A night watchman who dozes between rounds still writes the same number of entries in the book. The log looks complete. The hours it covers are not the hours you think.

The second thing is not fixed, and this log exists mostly to say so honestly.

My phone cannot reach the house. What does that mean? It means one of two things, and from where the phone is standing they are identical. Either the house is genuinely down, or my phone has lost its own connection and everything is fine without me. One vantage point cannot tell these apart. It is not a harder version of the same problem, it is a different problem, and no amount of cleverness inside the phone resolves it, because the phone is the thing in doubt.

What I did instead was stop pretending. The watchdog used to report that situation as "the network is down", which was a claim it could not support. Now it has a third state that says, in effect, "the house is unreachable from here and I have no second opinion". That is less satisfying and considerably more true, and the difference matters because an alert that overstates its certainty trains you to ignore alerts.

The real answer is a second vantage point: something outside my building that also checks, so two observers can disagree and the disagreement carries information. If the outside witness can reach the house and I cannot, the problem is me. I have not finished building that. The candidates are a small rented machine somewhere else, a friend's spare device, or a hosted check from a service whose only job is this. Each has a tradeoff I have not resolved, which is the honest state of it today.

The Build

The deadman switch pattern, the sleep and wake lock pitfalls that make it lie on a phone, how to label severity honestly when you do not know, and why a second vantage point is not optional. The last part is a design, not a finished build, and it is marked as such.

1. Build the deadman the right way round

A naive monitor alerts when it notices a failure, which means a monitor that dies becomes permanently reassuring. Invert it: the watched thing proves it is alive on a rhythm, and silence is the alarm.

# the watcher writes a timestamp every N seconds
#   heartbeat: { at: <timestamp>, by: <which watcher> }

# something ELSE checks the age of that timestamp
#   age > 3N   -> the watcher itself is the problem

Three intervals, not one. A single missed beat is normal jitter, and alerting on it produces noise that you will silence, which costs you the whole mechanism.

2. Expect the operating system to postpone you, and take a wake lock

On a phone, any timer in a background process is a suggestion. Hold a wake lock around the work and release it immediately after.

# the shape: acquire, do the work, release, ALWAYS release
acquire_wake_lock()
try:
    do_the_check()
finally:
    release_wake_lock()

Two traps. A lock you forget to release is a battery complaint, so put the release in the cleanup path rather than after the happy case. And a partial wake lock keeps the processor awake but does not restore the radio, so a check that needs the network may still need the device to be reachable; test on a phone that has genuinely been idle for an hour, not one you just unlocked.

3. Measure the real interval before trusting it

Do not assume the fix worked. Record when checks actually happen and look at the gaps.

# log every beat with its real timestamp, then read the deltas
#   expected: N, N, N, N
#   doze:     N, N, 14N, N, 9N      <- bursts after long gaps

# also exclude the device from battery optimisation, and verify that
# the setting stuck, because some systems quietly re-apply it

4. Give "I do not know" its own state, and do not colour it like a failure

This is the part most monitors get wrong, including mine. Three outcomes, not two.

# states, with honest meanings
#   OK              : I reached it
#   DOWN            : I could not reach it AND a second observer agrees
#   UNKNOWN         : I could not reach it and I have no second opinion

# and make the code refuse to report DOWN without the witness
#   witness_available AND witness_reached_it  -> DOWN
#   otherwise                                -> UNKNOWN

Write that condition explicitly rather than letting an absent witness fall through to the failure branch. A missing observer is not evidence of anything, and code that treats "no data" as "bad" will cry wolf every time your own connection hiccups.

5. Check your own connectivity first, and say which leg failed

Before claiming anything about the target, establish whether you have a network at all. One extra request buys most of the ambiguity back.

# order matters
#   1. can I reach the wider internet at all?        no  -> I am offline, say so
#   2. can I reach the target?                       yes -> OK
#   3. internet yes, target no, witness says yes     -> target is DOWN
#   4. internet yes, target no, no witness           -> UNKNOWN

This still cannot distinguish a target that is down from a path between you and the target that is down. That residue is what the witness is for, and nothing else solves it.

6. The second vantage point, which I have not finished

Published as a design rather than a result, because that is what it is.

# requirements
#   outside your building and off your connection
#   checks the SAME target on its own schedule
#   publishes a result your primary watcher can read
#   fails independently, or it is not a second opinion

# candidates, with the tradeoff
#   small rented machine : most control, a monthly cost, one more box to patch
#   a friend's spare     : free, depends on someone else's power and goodwill
#   hosted uptime check  : least work, least control, another account to manage

The trap to avoid: a witness that shares your failure modes is decoration. If it rides your connection, or reports through a service your primary watcher also depends on, it will go dark at exactly the moment you need its opinion.

7. Keep the honest catch written down next to the code

Until the witness exists, the monitor has a known blind spot, and the worst outcome is forgetting that while trusting the output.

# watchdog.md
#   KNOWN LIMIT: single vantage point. UNKNOWN means unreachable-from-here,
#   not down. Do not act on UNKNOWN as if it were DOWN.
#   OPEN: second vantage point not yet built. See candidates above.

Related: prove the monitor isn't the thing that's broken is the same question one layer in, the worst alert is a green one is what happens when a monitor fails quietly, and a health check that does real work is how a watcher becomes the outage.