- Prevented removing and renaming of the root directory. This is done by
repurposing the INVAL error in lfsr_mtree_lookup to indicate the
found entry is the root.
The root entry has special behavior in almost every function, owing to
the fact it doesn't really have an mid/rid. So I think this is a
reasonable approach.
- Added support for lfsr_stat of the root directory.
- Fixed off-by-two in lfsr_dir_seek thanks to the "." and ".." entries.
Humorously there is a comment noting this but the code didn't
actually match the comment.
Unfortunately the powerloss testing risks being a big time sink.
Figuring out the best scale of powerloss testing during normal testing
is probably going to be a constant balancing act.
Mainly trying to match the tests over mkdir/rm, which seem to have a
good amount of coverage.
- Fixed issue where move's desination rid wasn't updated correctly if
the destination split.
- Prevented renaming into nonexistant directories.
- Fixed neighboring rid adjustment in rename (+1 not -1 silly).
- Fixed erronously updating the grm's rid during lfsr_fs_fixgrm. In the
"I can't believe this ever worked" category, it seems this usually
didn't cause issues since mid was often marked as removed, making the
erronously updated rid ignored.
Only simple tests right now, but the theory is sound.
This mainly required the addition of the fancy in-device move attribute,
which copies all tags associated with an rid from one rbyd to another in
a single transaction.
This is a carryover from the previous littlefs implementation, though it
is easier to implement here since it is effectively a range query on the
rbyd tree, which trees are really good at. This was intentional.
Oh and I suppose this also required implementing lfsr_rename, which has
a few corner cases to watch out for.
It is nice that both lfsr_remove and lfsr_rename can rely on
lfsr_fs_fixgrm to finish all of the removes, which wasn't previously
reasonable due to the overhead of deorphaning.
Hopefully third times the charm.
The previous solution pretty bluntly did not work outside of the
recursive remove case, because the moment we mark the rid as deleted,
the directory positions no longer get updates. It's not possible to
update the directory position because we don't know how it maps into our
mtree without a full seek from the dstart.
After staring at it a bit, I think this solution should work:
1. Instead of marking the mid/rid as removed when dropping an mdir, we
set the weight to zero and the trunk to zero, causing mdir lookups to
return NOENT without actually going to disk.
This is very important since later mdirs could be allocated on the
same block, and going to disk can result in a corrupted lookup.
2. Eagerly seek to the next mid/rid after every lfsr_dir_read call. This
puts us in a position where rid can be >= the current mdir weight
without issues, and avoids degenerate cases that may be caused by
recursive removes.
3. If we remove an opened dir, instead of marking the mdir as deleted,
move the rid to the next rid. If the mdir was dropped, this leaves us
with rid == mdir weight, and the mdir trunk == 0.
The rid == mdir weight also occurs when we are creating a new file, so
we have a bit of common behavior we can rely on. We just need to make
sure that mdir updates respect the rid == mdir weight situation.
4. On each lfsr_dir_read call, we do an mtree seek of zero. This just
serves to fix our mdir if our rid == mdir weight, without much
additional code (yay for code reuse).
The use of weight=0, trunk=0, for a dropped mdir here is key, and makes
me wonder if this is a better indicator of a dropped mdir than another
reserved mid value. This probably deserves some investigation later.
"Recursion" here just refers to the ability to remove entries in a
directory while iterating over it. This is very useful when you just
want a directory gone, and can be extended to a "true" recursive remove
straightforwardly. This mainly tests that mid/rid updates in opened
mdirs are correct.
To make this work, we need to update opened dirs differently than files,
since opened dirs do not get marked as removed when its rid is removed
and contain an additional position in the dir that needs to be updated.
To keep track of the different types, littlefs now contains 2
linked-lists for opened mdirs. Maybe these should be correctly typed,
but by hiding the specific types behind an array of mdir linked-lists,
we can more efficiently iterate over both lists when necessary.
We should probably compare this approach to the type-tagged approach in
the previous littlefs implementation, but I think the idea of an array
of type-hidden linked-lists just didn't come to me then. There was also
a bit more room in the mdir structs to hide a 1-bit type field. The mdir
structs here are getting pretty squeezed since they are used everywhere.
These mirror the lfsr_mkdir tests, but backwards.
It's interesting to note the rm powerloss testing is much slower than
mkdir powerloss testing. This is because the rm tests can make
significant backwards progress if power is lost (these tests both make
and remove dirs), but mkdir tests always make forward progress (by only
making dirs).
In theory this is pretty much the same as lfsr_mkdir, but backwards.
The main work was making the interactions between removing mids/rids and
the grm correct. This ends up meaning we just need to update the grm on
any mid/rid update the same way we update the list of opened mdirs.
On the plus side, it turned out to be possible to deduplicate the mdir
uninlining route a bit, by adding range argument to lfsr_mdir_commit_
and changing the write of the newly uninlined mtree/mdir to marking
mtree as dirty and then joining the common path.
This lets us move the pre-commit round of grm updates into a single
location in lfsr_mdir_commit, removing and extra function definition and
the related state marshalling while also simplifying the control-flow.
This also raises the question, can more lfsr_mdir_commit be deduplicated
more? Uninlining is a infrequent operation we don't really need to
optimize for.
---
Testing lfsr_remove also found a bug related to incorrect propagation of
when the mroot becomes "unerased" (when rbyd overflows). This raises the
concern that we're not propagating unerased-states very rigorously, and
unexpected errors may not allow the filesystem to resume.
This has never been in a very good place for littlefs, but would be
worth improving in the future.
Instead of iterating over a number of seeds in the test itself, the
seeds are now permuted as a part of normal test defines.
This lets each seed take advantage of other test features, mainly the
ability to test powerlosses heuristically.
This is probably how it should have been done in the first place, but
the permutation tests can't do this since the number of permutations
changes as the size of the test input changes. The test define system
can't handle that very well.
The tradeoffs here are:
- We can't do cross-fuzz checks, such as the balance checks in the rbyd
tests, though those really should be moved to benchmarks anyways.
- The large number of cheap fuzz permutations skews the total
permutation count, though I'm not sure this matters.
before: 3083 permutations (-Gnor)
after: 409893 permutations (-Gnor)
To help with this, added TEST_PL, which is set to true when powerloss
testing. This way tests can check for stronger conditions (no EEXIST)
when not powerloss testing.
With TEST_PL, there's really no reason every test in t5_dirs shouldn't
be reentrant, and this gives us a huge improvement of test coverage very
cheaply.
---
The increased test coverage caught a bug, which is that gstate wasn't
being consumed properly when mtree uninlining. Humorously, this went
unnoticed because the most common form of mtree uninlining, mdir splitting,
ended up incorrectly consuming the gstate twice, which canceled itself
out since the consume operation is basically just xor.
Also added support for printing dstarts to dbglfs.py, to help debugging.
The grm bugs were mostly issues with:
1. Not maintaining the on-disk grm state in RAM (lfs->grm) correctly,
this needs to be updated correctly after every commit or littlefs
gets a confused.
2. lfsr_fs_fixgrm got a bit confused when it was missed when changing
the no-rm encoding from 0 to -2. Added some inline functions to help
avoid this in the future.
3. Leaking information due to mixing fixed sized and variable sized
encodings of the grm delta in places. This is a bit tricky to write
an assert for as we don't parse the full grm when we see a no-rm grm.
This implementation is in theory correct, but of course, being untested,
who knows?
Though this does come with remounting added to all of the directory
tests. This effectively tests that all of the directory creation tests
we have so far maintain grm=0 after each unmount-mount cycle. Which is
valuable.
This bug was just overlooked in testing the mtree, fortunately dir
fuzzing found it. Though since this depends on neighboring mdirs, it
probably would have been found quicker with smaller block sizes. At the
moment I am only testing on NOR-liked geometry (4KiB blocks).
The fix is easy, we can use the difference in the mtree size to
determine if a split or drop happened in mdir commit, since at most one
of these can happen on any mdir commit.
Also added an explicit test for mid updates when splitting and dropping.
lfsr_stat is really a directory operation underneath, so it's good to
add to our testing while we are building up the dir tests.
It's interesting to note lfsr_stat and lfsr_dir_read are less
deduplicatable than their previous versions, since lfsr_stat can get
most of it's info from lfsr_mtree_pathlookup. Though there will probably
need to be some code sharing when we get to files with sizes.
- Checksum collisions
- Collisions with root did
- Collisions needing wraparound
- Possible leb128 encoding issues
Sure enough the last one caught an off-by-one error in our calculation
of the leb128 encoded size. I sort of expected a bug there, since it's
rather nuanced math, so it's good to have test coverage now.
The main issues:
- The addition of the root's dstart entry during lfsr_format throws off
our mtree tests. It's a bit of a hack, but for now I am just manually
deleting the root's dstart entry at the beginning of each tests.
It might be possible to make the mtree tests work around the root's
dstart, but it seems to cause problems for when exactly the mtree
splits.
- btree dnamelookup and mdir dnamelookup need different things from
the rbyd dnamelookup when the dname is not found. The btree lookup
needs the largest branch smaller than the dname, since this is the
"bucket" containing our dname, while the mdir dnamelookup needs
the id that _follows_ the id smaller than the dname, since insertion
causes all ids >= the inserting id to shift up.
The solution here is to make rbyd dnamelookup behave as expected by
btree dnamelookup. btree needs more info about the branch (weight
mostly), so this avoids more issues. mdir dnamelookup adjusts the
id as needed, which costs a bit of code, but makes things work.
Fortunately, mdir dnamelookup can assume the weight is 1, which
simplifies things a bit.
This makes it now possible to create directories in the new system.
The new system now uses a single global "mtree" to store all metadata
entries in the filesystem. In this system, a directory is simply a range
of metadata entries. This has a number of benefits, but does come with
its own problems:
1. We need to indicate which directory each file belongs to. To do this
the file's name entry has been changed to a tuple of leb128-encoded
directory-id + actual file name:
01 66 69 6c 65 2e 74 78 74 .file.txt
^ '----------+----------'
'------------|------------ leb128 directory-id
'------------ ascii/utf8 name
If we include the directory-id as part of filename comparison, files
should naturally be next to other files in the same directory.
2. We need a way allocate directory-ids for new directories. This turns
out to be a bit more tricky than I expected.
We can't use any mid/bid/rid inherent to the mtree, because these
change on any file creation/deletion. And since we commit the did
into the tree, that's not acceptable.
Initially I though you could just find the largest did and increment,
but this gives you no way to reclaim deleted dids. And sure, deleted
dids have no storage consumption, but eventually you will overflow
the did integer. Since this can suddenly happen in a filesystem
that's been in a steady-state for years, that's pretty unnacceptable.
One solution is to do a simple linear search over the mtree for an
unused did. But with a runtime of O(n^2 log(n)), this raises
performance concerns.
Sidenote: It's interesting to note that the Linux kernel's allocation
of process-ids, a very similar problem, is surprisingly complex and
relies on a radix-tree of bitmaps (struct idr). This suggests I'm not
missing an obvious solution somewhere.
The solution I settled on here is to instead treat the set of dids as
a sort of hash table:
1. Hash the full directory path into a did.
2. Perform a linear search until we have no collision.
leb128(truncate28(crc32c("dir")))
.--------'
v
9e cd c8 30 66 69 6c 65 2e 74 78 74 ...0file.txt
'----+----' '----------+----------'
'-----------------|------------ leb128 directory-id
'------------ ascii/utf8 name
Worst case, this can still exhibit the worst case O(n^2 log(n))
performance when we are close to full dids. However that seems
unlikely to happen in practice, since we don't truncate our hashes,
unlike normal hash tables. An additional 32-bit word for each file
is a small price to pay for a low-chance of collisions.
In the current implementation, I do truncate the hash to 28-bits.
Since we encode the hash with leb128, and hashes are statistically
random, this gives us better usage of the leb128 encoding. However
it does limit a 32-bit littlefs to 256 Mi directories.
Maybe this should be a configurable limit in the future.
But that highlights another benefit of this scheme. It's easy to
change in the future without disk changes.
3. We need a way to know if a directory-id is allocated, even if the
directory is empty.
For this we just introduce a new tag: LFSR_TAG_DSTART, which
is an empty file entry that indicates the directory at the given did
in the mtree is allocated.
To create/delete these atomically with the reference in our parent
directory, we can use the GRM system for atomic renames.
Note this isn't implemented yet.
This is also the first time we finally get around to testing all of the
dname lookup functions, so this did find a few bugs, mostly around
reporting the root correctly.