Commit Graph

10 Commits

Author SHA1 Message Date
Christopher Haster 2835b17d14 Attempted to merge the mid's bid and rid into a single integer
This didn't really work out as well as I had hoped. There were a few
ideas on how to encode the bid/rid tuple without sacrificing the
(currently 31-bit) integer limit, but these just introduced too much
complexity.

Ideas:

1. In theory, as the mdirs increase in size, the quantity of mdirs needed
   for a given number of files decreases. If we say the number of files
   fits in an integer of a given size, than we can model the mapping to
   mdirs and rids roughly as the number of bits in that integer split
   between the two.

   Since the block_size is known, the we can find a rather conservative,
   yet useful, estimate of the upper bound of rids, which ends up
   being ~16 bytes ((2 alts + 1 null + 1 tag) * 4 bytes).

   And since our btrees are perfectly balanced, this encoding should only
   waste 1 or 2 bits due to rounding to rounding and sign encoding for
   special values.

     bbbbbbbb bbbbbbbb bbbbbbbr rrrrrrrr
     '-----------+-----------''----+---'
                 |                 '-- log2(block_size/32)-bit rid
                 '-------------------- remaining-bit bid

   Unfortunately, while this works ok on paper, and maximize the use of
   the bits we have available for the mid, the implementation ended up
   awkward and difficult to use.

   We need to either calculate the relatively complciated log2 of the
   block_size on the fly, or cache the value, and use it to shift the
   mid around to extract the bid/rid when needed.

   Unfortunately, perhaps due to the it being easy to use the bid/rid
   directly, we use and mutate the bid/rid quite a bit. We mutate when
   updating the mdirs, when decoding grms, when seeking mdirs, etc. If
   anything, updating the mid in total is rarer than updating the
   bid/rid component in complicated situations.

   Note to mention this required access to the lfs config to even begin
   decoding, complicating the API and making the result less efficient.

   Initial (unoptimized, and not even tested) code size showed ~+800
   bytes. So I decided to scrap this.

   Maybe it will be worth investigating dynamic rid sizes later, to
   increase the possible mtree size for a given mid width. Not sure.

2. Probably one of the worst ideas I've had so far, but it would solve
   the mid encoding problem, is to use some form a floating point to
   encode the bid/rid pair:

                          .----------.
                          v         .+-.
     bbbbbbbb bbbbbbbb bbbrrrrr rrrrssss
     '-----------+-------''----+---''-+'
                 |             |      '-- rid bits
                 |             '--------- variable rid
                 '----------------------- variable bid

    An even worse idea would be to use IEEE floating point here. Yes it
    would work, and probably work annoyingly well, but we it risk
    bringing in a lot of standard conforming backbending that we really
    don't care about.

    The idea here is to sacrifice some bits to encode the ratio of rid
    bits to bid bits. The value of this over the using the block_size is
    that we can decode the bid and rid using all of the bits in the
    integer alone. Avoiding memory access (and worse debugging) to load
    any external constants.

    As a plus, all mids in the system would have the same exponent,
    simplifying comparisons and other operations.

    But this is just trying to solve complexity by adding more
    complexity, so I'm not even going to try implementing it.

    Still, it's an interesting idea...

In the end I've gone with the KISS implementation. Use half-width
integers, in this case uint16s, for both the bid and rid:

  bbbbbbbb bbbbbbbb rrrrrrrr rrrrrrrr
  '-------+-------' '-------+-------'
          |                 '-- 16-bit rid
          '-------------------- 16-bit bid

This suffers from weakened limits around the number of rids in a block
and number of mdirs in the mtree, which is unfortunate. Still it is
probably worth the tradeoff for the RAM savings and encoding simplicity.

If the mdir is reasonably sized, this does probably approach a decent
distribution of rids and bids in 32-bits. But for outlier cases with
very small and very large mdirs, it risks premature out of bounds
errors.

To protect against mtree errors, we will probably need an additional
configuration option in the form of an mdir limit. Conveniently this
would also provide a way to enforce 2-block mode.

rid errors, on the other hand, depend on block_size/32, so we may not
need another configuration option and can rely on the block_size
to determine if the rids can overflow.

This is probably worth revisiting in the future. Fortunately, with
mdir_limit and block_size configuration options, it should be possible
to increase these limits in the future if this mid bid/rid design
changes.

            code          stack
  before:  22126           2136
  after:   22326 (+0.9%)   2088 (-2.2%)

This code size increase was unexpected. Maybe non-32-bit-aligned integers
cost more to load in thumb? Unsure.
2023-08-03 09:30:58 -05:00
Christopher Haster 5bdb55abec Fiddled with how opened mdirs are tracked and updated
The main intention here was to make the tracking of opened mdirs,
mostly opened lfsr_dir_t structs, simpler and more resilient to weird
corner cases. I'm not entirely sure this was successful.

The main changes:

- lfsr_dir_t now contains a full mdir for the dstart entry.

  This makes it so that dstarts are not a special case when it comes
  to mdir updates, though the fact that directories have 2 mdirs is
  still an awkward case on its own.

  I considered using two entries in the opened linked-list for this, but
  it wouldn't have worked out that well. Both entries need to update the
  directory position, so it would have required a third file type. We
  would also have needed to make sure removed mdirs mark both mdirs as
  removed, otherwise the position mdir would move around arbitrary into
  possibly erronous values.

  Instead the current solution treats the directory mdirs as a small
  array of 2 mdirs, which is as hacky as it is hacky, but does get the
  job done with little code duplication.

- Directory positions are updated a bit more intellegently.

  Instead of checking if in range before updating, which requires access
  to both mdirs and duplicate mid/rid comparison logic, position is
  updated without regard for the beginning of the directory, and
  un-updated if it was actually out of range of the directory.

  This means we only need to compare the mids/rids for each mdir once.

This changes make it so that lfsr_dir_rewind is much cheaper, and
doesn't even need to go to disk. Though I'm not sure it's worth the RAM
increase...

Expanding the lfsr_dir_t dstart entry to a full mdir does a lot for
making mdir updates more consistent, but increases the lfsr_dir_t size
from 52 bytes to 76 bytes (+46.2%).
2023-08-01 23:40:25 -05:00
Christopher Haster d8d8d1e2ac Dropped special LFSR_MID_RM mid
This is mostly to make it easier to merge mids/rids. Having a special
constant here is tricky when the mid/rid split point is dynamic.

Currently using rbyd.trunk=0 to indicate when an mdir is dropped. This
is nice as it preserves the last mid/rid, which is needed by the readdir
code, and it implicitly returns NOENT to all queries in
lfsr_rbyd_lookup.
2023-08-01 12:48:45 -05:00
Christopher Haster 18e1eb0b41 Moved rid into the mdir struct
When updating any opened mdirs to keep things in sync, we need to know
what rid the mdir is targeting in order to know which on-disk mdir it
should follow in the case of splits. Making this rid an actual member of
the mdir struct simplifies things.

This adds some RAM cost, though the plan is to merge the mid/rid into a
single integer, which requires this change and should actually save RAM
in the long run.

            code          stack
  before:  22342
  after:   22204 (-0.6%)   2144 (+1.1%)
2023-07-31 18:19:58 -05:00
Christopher Haster 9d0edea7e3 Reworked lfsr_rbyd_estimate to be a bit simpler
Instead of reading eagerly and retreating with the hopes of terminating
early (which almost never happens when compacting, since we need to find
the split_id). lfsr_rbyd_estimate now works inward from the first and
last id to find both the dsize and split_id.

One thing that helps this is the addition of a separate per-id
lfsr_rbyd_estimate, which will be useful for checking if the quantity of
file attributes overflows our mdir limitations.

lfsr_rbyd_estimate also now ignores the -1 id for split_id calculation,
since -1 ids are always cleaned up during splitting, though it does
include it in the calculated dsize so that the condition to split is
determined correctly.

---

This also required rebalance changes. Fortunately, one improvement here
is that we can make a simplifying assumption tha the number of tags
can't exceed the maximum possible number of tags in the calculated
dsize. So worst case, if every tag is empty, the maximum possible dsize
becomes 4*(2*log2(dsize/4))+dsize.

Though it's still unclear if rebalance is worth keeping. Current
comparison:
                  code          stack
  rebalance:     22362           2120
  no_rebalance:  21922 (-2.0%)   2120 (+0.0%)
2023-07-30 17:12:33 -05:00
Christopher Haster b1187595d6 Added support for recursive removes in directories
"Recursion" here just refers to the ability to remove entries in a
directory while iterating over it. This is very useful when you just
want a directory gone, and can be extended to a "true" recursive remove
straightforwardly. This mainly tests that mid/rid updates in opened
mdirs are correct.

To make this work, we need to update opened dirs differently than files,
since opened dirs do not get marked as removed when its rid is removed
and contain an additional position in the dir that needs to be updated.

To keep track of the different types, littlefs now contains 2
linked-lists for opened mdirs. Maybe these should be correctly typed,
but by hiding the specific types behind an array of mdir linked-lists,
we can more efficiently iterate over both lists when necessary.

We should probably compare this approach to the type-tagged approach in
the previous littlefs implementation, but I think the idea of an array
of type-hidden linked-lists just didn't come to me then. There was also
a bit more room in the mdir structs to hide a 1-bit type field. The mdir
structs here are getting pretty squeezed since they are used everywhere.
2023-07-25 12:54:49 -05:00
Christopher Haster c5e84e874f Changed how fuzz tests are iterated to allow powerloss-fuzz testing
Instead of iterating over a number of seeds in the test itself, the
seeds are now permuted as a part of normal test defines.

This lets each seed take advantage of other test features, mainly the
ability to test powerlosses heuristically.

This is probably how it should have been done in the first place, but
the permutation tests can't do this since the number of permutations
changes as the size of the test input changes. The test define system
can't handle that very well.

The tradeoffs here are:

- We can't do cross-fuzz checks, such as the balance checks in the rbyd
  tests, though those really should be moved to benchmarks anyways.

- The large number of cheap fuzz permutations skews the total
  permutation count, though I'm not sure this matters.

  before: 3083 permutations (-Gnor)
  after: 409893 permutations (-Gnor)
2023-07-18 21:40:44 -05:00
Christopher Haster cb1319c9e6 Expanded fuzz testing a bit, found/fixed an mid neighbor update bug
This bug was just overlooked in testing the mtree, fortunately dir
fuzzing found it. Though since this depends on neighboring mdirs, it
probably would have been found quicker with smaller block sizes. At the
moment I am only testing on NOR-liked geometry (4KiB blocks).

The fix is easy, we can use the difference in the mtree size to
determine if a split or drop happened in mdir commit, since at most one
of these can happen on any mdir commit.

Also added an explicit test for mid updates when splitting and dropping.
2023-07-07 16:31:52 -05:00
Christopher Haster 21f7fd1032 Added more testing over mkdir, fixed issues dname changes introduced
The main issues:

- The addition of the root's dstart entry during lfsr_format throws off
  our mtree tests. It's a bit of a hack, but for now I am just manually
  deleting the root's dstart entry at the beginning of each tests.

  It might be possible to make the mtree tests work around the root's
  dstart, but it seems to cause problems for when exactly the mtree
  splits.

- btree dnamelookup and mdir dnamelookup need different things from
  the rbyd dnamelookup when the dname is not found. The btree lookup
  needs the largest branch smaller than the dname, since this is the
  "bucket" containing our dname, while the mdir dnamelookup needs
  the id that _follows_ the id smaller than the dname, since insertion
  causes all ids >= the inserting id to shift up.

  The solution here is to make rbyd dnamelookup behave as expected by
  btree dnamelookup. btree needs more info about the branch (weight
  mostly), so this avoids more issues. mdir dnamelookup adjusts the
  id as needed, which costs a bit of code, but makes things work.

  Fortunately, mdir dnamelookup can assume the weight is 1, which
  simplifies things a bit.
2023-07-06 00:55:28 -05:00
Christopher Haster 2fe2078f50 Renamed tests/benches such that order is logical
It doesn't make sense to test more complex logic, such as t2_btree.toml,
when the logic it is built on, t1_rbyd.toml, does not past testing. The
test runner already guarantees a consistent lexicographic order, so all
we need to do is renamed these from test_* -> tn_*.

Note, if we every have more than 10 tests, we will need to bump up the
number of digits for all tests, so t1_rbyd.toml -> t01_rbyd.toml. This
is the main downside of lexicographic ordering. But we'll cross that
bridge when we get to it.
2023-06-30 16:37:23 -05:00