On 16 August we installed some wireless testing tools on the wrong
machine. Eight seconds later, airmon-ng check kill ran.
That command does exactly what it says, and it is documented: it stops the services managing your wireless connection so a card can enter monitor mode. It did. The connection dropped, nothing brought it back, and the machine had to be rebooted by hand. Four jobs that were running at the time died with it, along with the dashboard and the local model server.
The command was fine. The machine was wrong. Wireless testing belongs on a separate box, and we already had that written down.
The gate we built
This machine has two wireless cards. The built-in one carries the
connection. A USB adapter is the one meant for testing. The problem
with that command — and with rfkill block wifi, and
with stopping the network manager outright — is that none of
them can be aimed. They are global switches, so they always take
the card you are depending on.
So we added a gate that reads every shell command before it runs and refuses the ones that would drop the connection, naming the interface it is protecting. It resolves that interface by looking up which card is carrying the live network at that moment, rather than trusting a name written into the file months earlier. It denies outright rather than pausing to ask, because there is not always somebody there to answer.
It has 36 test cases, and the ones that matter most are the should-allow cases. Checking the status of the network manager passes. Putting the USB adapter into monitor mode passes. A gate that blocks ordinary work is a gate people learn to route around, and a gate people route around protects nothing.
Then it blocked the write-up
Filing a note about all this, through the shell, was refused. The note quotes the command. The gate matched the string in the text being written to a file, not a command being run.
It refused rather than allowed, which is the direction you want, and the workaround took seconds — write the file with something that writes files instead of going through a shell.
But the limit underneath it is real. Any note, runbook, brief or incident report that documents a dangerous command cannot be written that way, and those are precisely the documents an operation like ours produces most. Our entire discipline is writing down what went wrong. A filter that cannot tell the difference between running a command and describing one will fire hardest on exactly the material that stops the incident happening twice.
Why we did not just widen it
The fix worth making is to match on the command being invoked rather than on any appearance of the string in the payload. That is a change to a safety gate, and it gets made deliberately, with the test cases updated, on a day when that is the job.
Loosening a guard as a side effect of being mildly inconvenienced by it is how guards quietly stop existing. The inconvenience was thirty seconds. The thing the guard prevents cost us a reboot, four running jobs and somebody walking over to the machine.
The general shape
Every filter has a false positive rate, and where those false positives land is not random. A filter aimed at dangerous things will catch documentation about dangerous things, security notes, training material and postmortems — the writing that exists specifically to reduce the risk it is matching on.
Worth knowing before you conclude your guard is broken. Ours was not broken. It was too literal, in the safe direction, at the one moment that was inconvenient.
What to do
Take one guardrail you rely on and ask what it actually matches on: the action, or the text. Then try writing your own incident report through it. If that fails, you have found the limit before it finds you.