Kyber Cypher v008/v007/v006/v005/v004/v003 KC//NODE-01 00:00:00:00

The lock only half your writers take

Field Log // 031 Status Live Difficulty Free Cost Nothing
The Story

Two processes were writing to the same file. One trimmed it on a schedule, the other added to the top of it whenever I had something to record. Both read the whole file, changed it, and wrote it back. If they ever overlapped, the slower one would finish last and silently erase the other's work.

They overlapped. I watched a file grow by several kilobytes while a trim was running, which meant the two had been in the same window, and I got lucky on the ordering. Nothing was lost that time. That is not a safety property, that is a coin landing the right way up.

So I added a lock to the trimmer. Done, I thought, and then I said the sentence out loud: the trimmer takes a lock. And the other writer? The other writer was me, in a terminal, pasting text into a file. I cannot make myself take a lock. I cannot make a future script someone writes take one either.

A one way sign only works on drivers who read signs. If half the traffic is pedestrians who never look up, you have not made the street one way. You have made it one way for cars.

That is the honest shape of a lock: it binds the processes that take it, and it is completely transparent to the ones that do not. Adding it to one writer makes that writer polite. It does not make the file safe. Reporting it as "fixed" would have been a lie that only showed up later, when something was missing and I believed the problem had been solved months earlier.

So I reported it as partial, which is what it was, and then added a backstop for the writers I cannot bind. If the file was modified in the last few seconds, assume someone is mid-write and skip this cycle. That is not a guarantee. It is a reduction in the size of the window, which is the honest best available when you do not control every writer.

The other thing I learned is that locking and atomic writing are different problems and I had been treating them as one. A lock stops two writers from interleaving. It does nothing about a single writer dying halfway through and leaving a truncated file. You need both, and I only had the one I had thought of.

The Build

A file locking pattern with clean skip on contention, atomic write as the separate protection against a crash, a stability backstop for writers you cannot bind, and how to test the race on purpose rather than waiting for it.

1. Use the operating system's lock, not a file you create yourself

An advisory lock held by the kernel is released automatically when the process ends, including when it is killed outright. That single property removes a whole category of problem.

# Python
import fcntl
fh = open(LOCKFILE, "w")
fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB)   # non-blocking

# shell
flock -n <lockfile> -c '<your command>'

2. Do not write a PID file, and do not invent a staleness rule

This is the mistake I see most. People write their own lock as a file containing a process id, then need a rule for when a leftover one is stale, then that rule wedges the job forever or breaks a lock that was genuinely held.

# NOT needed with a real lock:
#   a PID file
#   an "if older than N minutes, break it" heuristic
#   a cleanup step on startup
# the kernel drops the lock when the holder's descriptor closes,
# including on a hard kill. A leftover lock FILE is harmless.

3. On contention, skip cleanly. Never wait forever

For a scheduled job, a missed cycle is nothing, because the schedule comes round again. A job that blocks indefinitely inside a timer is a hang you will debug at a bad time.

# scheduled work: non-blocking, skip and log
#   could not acquire -> "another writer holds it, leaving this cycle" -> exit 0

# interactive work: a short timeout is fine, then fail LOUDLY
#   waited 30s, still locked -> refuse, tell the human, change nothing

Different callers want different behaviour from the same lock, so make the timeout a parameter rather than a constant.

4. Scope the lock per resource, not globally

One lock for everything turns unrelated jobs into queue mates. Derive the lock name from the file so any writer can compute it without sharing code.

# derive it from the target's name
#   <lockdir>/.<resource-name>.lock
# so a second writer, in another language, can take the SAME lock
# without importing anything from you

5. Add atomic write, which is a different protection

Write to a temporary file in the same directory, flush it to disk, then rename over the target. A rename within one filesystem is atomic, so a reader sees either the whole old file or the whole new one, and a crash mid-write leaves the original intact.

fd, tmp = mkstemp(dir=same_directory_as_target)
write(fd, data)
flush(fd); fsync(fd)        # on disk, not just in a buffer
rename(tmp, target)         # atomic

Same directory matters. A rename across filesystems is a copy and a delete, which is not atomic and defeats the point.

6. Add a stability window for the writers you cannot bind

For anything that writes by hand or by a script you do not control, check how recently the file changed and stand down if it was just touched.

# before modifying, ask how old the last change is
age = now - modified_time(target)
if age < STABILITY_SECONDS:        # 50s is a reasonable start
    skip("recently modified, a write may be in flight")

# and re-check it again just before committing, because an unlocked
# writer may have landed while you were working
if modified_time(target) != mtime_when_i_read_it:
    abort_without_writing()

That second check is the one that saves you. It converts a silent overwrite into a skipped cycle.

7. Test the race deliberately

Do not wait to observe it in production. Create it, on copies, and confirm both sides survive.

# on throwaway fixtures, in this order
#   1. hold the lock in one process; run the job -> expect clean skip, no change
#   2. kill -9 the holder; run the job          -> expect it proceeds, not wedged
#   3. run job and a LOCKED writer at once      -> expect serialised, nothing lost
#   4. run job and an UNLOCKED writer at once   -> expect the backstop to abort
#   5. account for every line afterwards: is it in one place or the other?

Test four is the one that matters, because it is the case the lock does not cover.

8. Report a partial fix as partial

Write down which writers take the lock and which do not. A reader who believes the file is protected will later make a decision that assumes it.

# locking.md
#   takes the lock : the scheduled trimmer, the helper tool
#   does NOT       : anything written by hand, any future ad hoc script
#   those are covered only by the stability window, which is a
#   backstop and not a guarantee

The honest catch: two of three writer classes taking the lock is not a solved problem, it is a smaller problem. The remaining one is a human, and humans are not bindable by file descriptors.

Related: archive it, prove the archive has it, then delete is the other half of safely shrinking a file.