35db3bc97f
After thinking about this for a while, btree node compaction is
subtlety different from mdir compaction, less valuable, and adds more
risk:
- Unlike mdirs, btree node compaction will always allocate a new
block, leading to a higher chance of alloc failure.
- Btree node compaction also always requires additional writes to
propagate btree changes, whereas mdir compaction is usually
self-contained unless it triggers a relocation. If btree nodes are
mostly full this risks being counter-productive.
- Btree node compaction requires a full tree traversal, whereas mdir
compaction requires only traversing the mtree. Though you can always
force mtree-only traversal manually with LFS_GC_MTREEONLY.
- Btrees/bshrubs are also more likely to be "cold storage", that is it
probably won't be uncommon to create long-lived read-only btrees as a
part of files. Compacting these btrees can actually be counter-
productive as it can encourage splitting.
- Btrees/bshrubs are also more likely to be one use, and discarded as a
file is truncated and rewritten. Compacting btree nodes in this case
is a waste of erase cycles.
And since btree node compaction also introduces a lot of complexity/risk
of bugs, I'm going to drop this for now and limit LFS_GC_COMPACT to only
compacting mdirs. At least this tested implementation will live in the
history and can always be reintroduced in the future if it becomes a
wanted feature.
---
As is usually the case, doing less work ends up with less code:
code stack
before: 36292 2704
after: 35888 (-1.1%) 2696 (-0.3%)
Note this still keeps the rbyd-specific commit logic necessary for
committing to specific btree nodes, even though btree node compaction
was the only current use case. This should eventually be useful for
metadata repair. Hopefully const-propagation can minimize the cost, but
realistically this means we're probably leaving some code savings on the
table.