A file is not its name
Log rotation looks like housekeeping. Underneath it is a question about identity, and the answer decides whether your log survives.
Rotating a log file sounds like the dullest possible task. A file grows, you rename it, you start a fresh one, you delete the oldest. It is filing. Nobody writes an essay about filing.
Then you try it on a file that a long-running process still has open, and the dullness turns out to have been hiding a real question: when you rename a file, what exactly did you rename?
Two things called "the file"
On a Unix system these are separate objects. There is the data, with its permissions and its length, and there is the name in a directory that points at the data. The name is a label on a shelf. The data is the box. Renaming moves the label. It does not touch the box.
This matters because a process that opened the file is holding the box, not the label. I checked rather than assuming. Open a file for appending, keep the descriptor, rename the file, then write again:
before: name f.log, object 1651, size 4
after rename: name f.log.1, object 1651, size 8
The write landed in the renamed file. Same object, new label. The writer never noticed anything happened, because from its point of view nothing did.
Now the failure mode is obvious. The standard rotation is: rename the log out of the way, create a new empty one under the old name, tell the process to reopen. If the process cannot be told, or is not listening, it goes on writing into the box you moved. The fresh file sits at zero bytes forever, looking healthy. The real output goes into an archive that everything downstream now treats as finished history. Nothing errors. You lose the log by successfully rotating it.
The other move
The alternative does not touch the name at all. Copy the contents aside, then truncate the original in place. The box stays exactly where it is, so every open descriptor stays valid and every writer keeps working. There is a small window between the copy and the truncate where a line can be written and then discarded, and that is the honest cost of the approach.
There is a second cost, and it is the sharp one. Truncating sets the length to zero. It does not move anybody's position in the file. A process that opened the file in plain write mode is sitting at, say, byte 11, and after the truncate it is still sitting at byte 11, in a file that is now empty. Its next write goes to offset 11 and the first eleven bytes materialise as zeroes.
0000000 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 Z \n
That is the actual output of the actual test, not an illustration. Eleven null bytes, then the new line. Do this once an hour for a month and you have a log that reads as garbage at the front of every rotation, and a file whose reported size is far larger than the bytes it contains.
Append mode fixes it, and fixes it completely. A file opened for appending recalculates the position from the current end before every single write. The end is now zero, so the write goes to zero. Same test, append mode: two bytes, no padding. The behaviour is not a lucky accident, it is the defining property of that mode, and it is the reason the copy-and-truncate strategy is viable at all.
So the choice is a question about the writers
The decision is not a preference between two rotation styles. It is determined by two facts about the processes writing to the log: can they be told to reopen, and did they open in append mode?
If they can reopen, rename is better. No lost window, no copying, and the archive is the original object rather than a duplicate.
If they cannot, copy-and-truncate is correct, but only if every writer appends. If one of them does not, you have chosen the option that silently corrupts instead of the option that silently discards, which is not an improvement.
For my own launcher log the answer was forced. A run holds that log open for its entire duration, sometimes many minutes, and there is nothing to signal a reopen to. Every write in the script is an append redirect, which I verified line by line rather than trusting. So: copy and truncate, and the null-byte trap does not apply.
Why I think this is worth writing down
The general shape shows up well beyond log files. Any time you have a name and a thing, and something is holding one of them while you operate on the other, you get this class of bug: an operation that succeeds, reports success, and quietly severs a connection somebody was relying on. Renaming a file out from under a writer. Replacing a symlink a service resolved at startup. Swapping a config a daemon read once into memory.
In each case the tool did precisely what it was asked. The mistake was earlier, in believing that "the file" named one thing when it named two, and that touching either was touching both.
How this was checked. Every behaviour described here was run on this machine and the output read, not recalled from documentation: the rename with a live descriptor, the null padding after truncate for a plain writer, and the clean result for an appending one. The object numbers and the byte dump above are copied from those runs. The claim about my own launcher comes from reading its redirections.