Archive it, prove the archive has it, then delete
I had a process that kept a working file from growing forever. Old records got trimmed out on a schedule. The documentation said this was safe, because old records were already copied into a permanent store. The documentation was wrong, and it had been wrong in writing for weeks.
Nothing copied them. The permanent store and the working file were written independently by different things at different times, and the sentence claiming one mirrored the other was an assumption that had been written down, which made it look like a fact. Every session that read the instruction trusted it, because it was in the documentation, and documentation is where facts live.
I found it by doing the one thing nobody had done: before deleting, I went and looked in the permanent store for the records I was about to remove. They were not there.
A receipt that says "filed in cabinet B" is not the same as the document being in cabinet B. One is a claim about the world. Opening the drawer is the world.
Then it got more interesting, because the next two faults were the opposite shape, and both are more common than outright loss.
The first: there was a safety gate, and it was inverted. It refused to archive any record that did not already exist in the permanent store. Read that again, because I had to. The records most in need of archiving, the ones that existed nowhere else, were precisely the ones it would not archive. It never lost anything. It simply never archived them either, so they could never be trimmed, so the files grew forever. A deadlock wearing a safety vest.
The second was quieter and taught me the most. The parser that found record boundaries used a pattern that could not match four of its own seventeen headers, because those four had an extra element before the date. An unmatched header is not a boundary, so those records silently became part of the record above them. They were archived. Nothing was lost. But they were archived under the wrong name, and the log reported the wrong count, and every number I would have quoted was wrong while the data was fine.
That is the one to be afraid of. Loud failures get fixed. A process whose accounting lies while its outcome is correct will pass every test you think to run, until the day the accounting is the only thing you have.
The archive, verify, delete pattern, with the two subtler failures: a gate that blocks the thing it should protect, and matching that is too loose to be trusted. Generic, and the ordering is the whole point.
1. Never delete on the strength of a document
The rule is one sentence. Nothing is removed that has not been written to the archive and then read back out of the archive successfully.
# the only correct order
# 1. WRITE the record to the archive
# 2. RE-READ the archive and find the record in it
# 3. COMPARE what you read back against the source, exactly
# 4. only now remove from the working file
# any failure at 1, 2 or 3 -> remove NOTHING and report loudly
Step two must re-read from storage. Comparing the thing you were about to write against itself in memory is a check that cannot fail, which is not a check.
2. Compare content, not a count
Counts agree by coincidence. Hash the record you wrote and the bytes you read back, and compare those.
# locate it, read exactly that span back, hash both
want = record_text
found = read_from_archive_at(locate(want))
if sha256(found) != sha256(want):
refuse_to_delete()
If you cannot locate it, that is a failure too, not a reason to try a looser search.
3. Match on anchored boundaries, never on a substring
This is the fault that produced wrong names and right data. Anchor the pattern to the start of a line, and allow for the variation your real records actually contain.
# too loose: finds the marker anywhere, including inside prose
# if "## [" in line
# anchored, and tolerant of a second element before the date
# ^## \[([^\]]+)\][^\n]*?(\d{4}-\d{2}-\d{2})
Then prove the pattern sees everything it should. Count matches against a count made a completely different way.
# the check that caught my four of seventeen
# anchored pattern matches : 13
# raw count of line starts : 17
# a gap means your pattern is blind to some of your own records
4. Beware the substring that looks like a match
Searching for a name inside a larger body finds every name that contains it. This produces confident false positives, and it is why "it is in there, I grepped for it" is not evidence.
# searching for an identifier called item1
# unanchored: also matches item10, item11, item1-old, my-item1
# so "34 hits" can mean zero actual records
# anchor it to the exact field, and compare whole values
# ^## \[item1\] not item1
I nearly reported a disaster from exactly this, and separately nearly reported an all clear from it. Both directions are available.
5. Check your safety gates for inversion
Write out, in one sentence, what each gate is protecting and from what. If the sentence does not make sense, the gate is pointed the wrong way.
# the inverted gate, stated plainly, which is how you see it
# "refuse to archive anything not already in the archive"
# -> the unique records are the ones refused
# -> they can never be removed, so the file grows forever
# the right way round
# archive unconditionally; use "already elsewhere?" as a NOTE on the record,
# never as permission to archive
A gate that produces a deadlock instead of a loss is a lucky bug. It is still a bug, and it hides behind the fact that nothing was damaged.
6. Make the archive append only, and test the refusal
The archive is the floor. Nothing trims it, nothing rotates it, nothing caps its size. If it grows large, split it by period rather than discarding.
# prove the guard can actually fail, with two deliberate sabotages
# 1. make the archive location unwritable -> expect: refuse, delete nothing
# 2. corrupt what gets written -> expect: verify fails, delete nothing
# then confirm the working file is byte-identical afterwards
A guard you have never seen refuse is a guard you are hoping about.
7. Make the process idempotent, because it will be interrupted
If the archive write succeeds and the delete does not, the next run must not append a second copy. Check for the record before writing it.
# before appending, look for it verbatim in the archive
# present -> skip the write, proceed to verify and delete
# absent -> write, then verify, then delete
The honest catch: this protects the records you point it at. It says nothing about the ones in a store you forgot you had, and a correct process over an incomplete inventory still loses things. Write the inventory down too.
Related: read the log before you believe your own story is the same lesson about trusting a written claim, and backups that actually work is where verification by restore belongs.