Commit Graph

18 Commits

Author SHA1 Message Date
Christopher Haster db1f941e90 Slightly reworked lfs3_file_opencfg's mid reservation path
And tried to more consistently use lfs3_path_namelen.

In a perfect world we would just use lfs3_path_namelen everywhere and
let the compiler figure it out, but unfortunately this leads to poor
code generation in some places, even with __attribute__((pure)) hacks.

Code changes:

           code          stack          ctx
  before: 37832           2416          636
  after:  37824 (-0.0%)   2416 (+0.0%)  636 (+0.0%)
2025-06-24 15:16:55 -05:00
Christopher Haster 1b76bd04ce kv: Some minor file cache_buffer tweaks
- Unconditionally pass buffer as cache_buffer in lfs3_set now that we
  rely on LFS3_o_WRSET

- Swapped true -> 1 for non-null don't-care buffer pointer

Saved one instruction as expected for the conditional assignment, but
added a bit of stack. Weird, but probably just compiler noise:

           code          stack          ctx
  before: 37836           2408          636
  after:  37832 (-0.0%)   2416 (+0.3%)  636 (+0.0%)
2025-06-22 15:55:14 -05:00
Christopher Haster e7c7a81cfe Revisited zero-length file sync path
This needed a second pass. Changes:

- Small file flushes are no longer limited to LFS3_o_UNFLUSH, which
  should avoid bshrubs/btrees being written for small files with
  complicated seek+writes. Now, any file small enough is converted
  to a small file when we would need to flush.

  This does _not_ flush small unsync files that don't need to be
  flushed, though I'm not exactly sure how that would happen (broadcast
  from file with a different cache size?)

  I think this was a regression from previous logic.

- discardbshrub/discardbleaf moved into lfs3_file_sync_, otherwise
  we risk discarding the bshrub/bleaf without setting UNSYNC.

  This keeps all the state changing logic together.

- We now use lfs3_file_size_ == 0 as the decision for committing bnulls.

  size_ == 0 implies bnull, and this avoids the extra headache of
  checking for pending small file flush.

Note the ultimate decision on if the file is small is still left up to
lfs3_file_sync. lfs3_file_sync_ just relies on the UNFLUSH + UNCRYST +
UNGRAFT checks to do the last minute small file flush (aside from
asserts).

The UNFLUSH + UNCRYST + UNGRAFT checks look a bit messy, but keep in
mind these optimize to a single bitmask.

Saves a tiny bit of code:

           code          stack          ctx
  before: 37856           2416          636
  after:  37836 (-0.1%)   2408 (-0.3%)  636 (+0.0%)
2025-06-22 15:37:53 -05:00
Christopher Haster 7a6aad3cc8 Cleaned up potential lfs3_mdir_commit dedup TODOs
Unfortunately neither of these were actually deduplicatable:

1. We can't easily move dir update logic into lfs3_mdir_commit, because
   lfs3_mdir_commit has no knowledge of the current did.

   Maybe we can add did-related nudge functions, but the logic would
   still need to be external to lfs3_mdir_commit. lfs3_mdir_commit only
   understands mids.

2. lfs3_alloc_ckpoint continues to be enticing, but fortunately a
   previous commit reminded me that we explicitly need to _not_ call
   lfs3_alloc_ckpoint before the lfs3_mdir_commit in
   lfs3_bshrub_commitroot_.

   In theory we could add lfs3_mdir_commit and lfs3_mdir_commit_ to
   make lfs3_alloc_ckpoint opt-out, but the lfs3_mdir_commit is already
   a bit of a mess. And maybe keeping the lfs3_alloc_ckpoint calls
   explicit is a good thing. It's better to ENOSPC than double alloc a
   block.
2025-06-22 15:37:47 -05:00
Christopher Haster f967cad907 kv: Adopted LFS3_o_WRSET for better key-value API integration
This adds LFS3_o_WRSET as an internal-only 3rd file open mode (I knew
that missing open mode would come in handy) that has some _very_
interesting behavior:

- Do _not_ clear the configured file cache. The file cache is prefilled
  with the file's data.

- If the file does _not_ exist and is small, create it immediately in
  lfs3_file_open using the provided file cache.

- If the file _does_ exist or is not small, do nothing and open the file
  normally. lfs3_file_close/sync can do the rest of the work in one
  commit.

This makes it possible to implement one-commit lfs3_set on top of the
file APIs with minimal code impact:

- All of the metadata commit logic can be handled by lfs3_file_sync_, we
  just call lfs3_file_sync_ with the found did+name in lfs3_file_opencfg
  when WRSET.

- The invariant that lfs3_file_opencfg always reserves an mid remains
  intact, since we go ahead and write the full file if necessary,
  minimizing the impact on lfs3_file_opencfg's internals.

This claws back most of the code cost of the one-commit key-value API:

              code          stack          ctx
  before:    38232           2400          636
  after:     37856 (-1.0%)   2416 (+0.7%)  636 (+0.0%)

  before kv: 37352           2280          636
  after kv:  37856 (+1.3%)   2416 (+6.0%)  636 (+0.0%)

---

I'm quite happy how this turned out. I was worried there for a bit the
key-value API was going to end up an ugly wart for the internals, but
with LFS3_o_WRSET this integrates quite nicely.

It also raises a really interesting question, should LFS3_o_WRSET be
exposed to users?

For now I'm going to play it safe and say no. While potentially useful,
it's still a pretty unintuitive API.

Another thing worth mentioning is that this does have a negative impact
on compile-time gc. Duplication adds code cost when viewing the system
as a whole, but tighter integration can backfire if the user never calls
half the APIs.

Oh well, compile-time opt-out is always an option in the future, and
users seem to care more about pre-linked measurements, probably because
it's an easier thing to find. Still, it's funny how measuring code can
have a negative impact on code. Something something Goodhart's law.
2025-06-22 15:37:07 -05:00
Christopher Haster 0772d10dbc kv: Implemented one-commit lfs3_set
This reworks lfs3_set to be able to write small files in a single
commit, by duplicating most of lfs3_file_opencfg.

The only real issue with the naive key-value API was the forced double
commit in lfs3_set. It may not seem like much, but on storage with large
prog sizes (NAND), the difference can be significant.

How significant? Well the difference approaches ~2x. Not because of the
inherent cost of progs, but because prog alignment will force you to
erase ~2x as often.

This small file logic matches lfs3_file_sync's small file logic, so if
you can lfs3_file_sync in one commit, you should be able to lfs3_set in
one commit.

It actually just uses lfs3_file_sync's small file logic for _existing_
files, but unfortunately we need special handling for _non-existing_
files to avoid the stickynote in lfs3_file_opencfg. Fortunately the
small file shrub commit is not too tricky to create on-demand. And as a
funny coincidence, _non-existing_ files, by definition, can't have any
opened file handles, so we don't need to worry about the missing file
broadcast logic.

---

Unfortunately, it turns out duplicating most of lfs3_file_opencfg adds a
huge chunk of code:

              code          stack          ctx
  before:    37644           2448          636
  after:     38232 (+1.6%)   2400 (-2.0%)  636 (+0.0%)

  before kv: 37352           2280          636
  after kv:  38232 (+2.4%)   2400 (+5.3%)  636 (+0.0%)

So may need to go back to the drawing board.
2025-06-22 15:36:36 -05:00
Christopher Haster a75537faff kv: Implemented a simple key-value API
This adds a couple functions that treat files as simple key-value pairs:

- lfs3_get    - Read a file
- lfs3_size   - Get the size of a file
- lfs3_set    - Write a file
- lfs3_remove - Remove a file (this one already exists!)

The idea is the only real difference between a filesystem and key-value
store in the microcontroller space is the API, and the key-value API
_is_ much easier to use.

It also opens the door to making the file API opt-out in the future to
trade code cost for feature set. littlefs will probably never be
competitive with other microcontroller-scale key-value stores, but it
may be interesting for systems already using littlefs for other storage.

And don't worry, these are still files, so they can always be opened
with the full file API when more advanced operations are needed.

These APIs also matches the custom attribute APIs, which makes sense
because they're both key-values. Any mismatch should be considered an
API bug, because the best user interface is a consistent one.

This new API is tested in tests/test_kv.toml.

---

At the moment the implementation is naive, just sitting on top of the
file API. This works remarkably well thanks to littlefs's cache
bypassing logic, but does have some downsides:

- lfs3_set always writes two commits: one for the stickynote and one for
  the file sync.

  Unfortunately this is a fundamental limitation of littlefs's file API.
  One nice benefit of lfs3_set is in theory we can bypass this
  limitation, but not if we just sit on top of the file API.

- There may be code savings from more tightly integrating the key-value
  code.

This also highlighted an awkward corner case with per-file cache
configuration in which the buffer needs to be non-null even if zero. Not
the end of the world, but just a bit awkward. Maybe this deserves
revisiting in the config API rework?

---

Code changes were relatively minimal given that this is a whole new API,
unfortunately the stack took quite a hit:

           code          stack          ctx
  before: 37352           2280          636
  after:  37644 (+0.8%)   2448 (+7.4%)  636 (+0.0%)

The stack surprised me, but in hindsight it makes sense. In sitting on
top of the reset of the codebase, the key-value API adds very little
code, but every stack allocation in these functions add to the stack
hot-path.

This isn't the end of the world, and it's actually probably a good thing
to have an lfs3_file_t allocated in the stack hot-path. lfs3_file_t's
size has been a bit difficult to track thanks to struct lfs3_info
dominating ctx measurements...
2025-06-22 15:22:21 -05:00
Christopher Haster 40a8c02604 Moved ifdefs after comments
So:

  // blablabla this is my cool function
  #ifdef LFS3_COOL
  int lfs3_cool(lfs3_t *lfs3);
  #endif

Mainly because this reads better and moves the compilation conditions
closer to the actual declaration.

One concern is if this will interfere with future doxygen/documentation
generation, but I think we can expect future scripts to be able to parse
relevant ifdefs. For one, we want to make sure to include any required
ifdefs in generated documentation, so if a script can't even parse
ifdefs, uhhhhh...

No code changes.
2025-06-06 01:43:58 -05:00
Christopher Haster b5568d076b Dropped LFS3_DATA_GRM
I think this was just missed during the various lfsr_data_t reworks.

No code changes.
2025-06-06 01:18:18 -05:00
Christopher Haster 0096305968 rdonly: Dropped rbyd.eoff when LFS3_RDONLY
rbyd.eoff has the relatively unique property of only being useful in
rdwr mode. In rdonly mode we don't care where the next erased-state
starts because we're never going to use it.

Since rbyds are used everywhere, dropping rbyd.eoff has the potential to
save a significant amount of RAM.

---

At least on paper. We were using the field in lfs3_rbyd_fetch to keep
track of the most recent valid commit perturb/eoff, which was a bit
tricky to disentangle.

Disentangling lfs3_rbyd_fetch does add a bit of code to the default
build, but saves code, stack, and ctx in the rdonly mode:

                   code          stack          ctx
  rdonly before:  10640            816          524
  rdonly after:   10616 (-0.2%)    808 (-1.0%)  508 (-3.1%)

  default before: 37320           2280          636
  default after:  37352 (+0.1%)   2280 (+0.0%)  636 (+0.0%)

In theory we could ifdef the crap out of lfs3_rbyd_fetch to claw back
this code, but 1. 32 bytes of code is really not that much code, 2. the
more rdonly and default diverge the more likely rdonly breaks, and 3. I
think the new code is a bit more readable since it avoids masking
perturb/eoff together until the last minute.
2025-06-06 01:17:12 -05:00
Christopher Haster 7cc87a4fe6 rdonly: Dropped file.b.shrub_ when LFS3_RDONLY
We don't need the staging shrub if we never stage shrubs!

The only hangup was reuse of the staging shrub to load bshrubs/btrees in
lfs3_file_fetch (we need to be able to fallback to the previous shrub if
we error in lfs3_file_resync), but this can be handled with a stack
allocated shrub.

If btree-leaf-caches make a return, we would need to stack allocate this
anyways due to the lopsided cost of the main/staging btrees/bshrubs
introduced to avoid wasting space on the useless
staging-shrub-leaf-cache.

This saves some code in LFS3_RDONLY, and apparently an instruction or
two in the default build (I guess stack loads/stores are cheaper?):

                   code          stack          ctx
  rdonly before:  10680            840          524
  rdonly after:   10640 (-0.4%)    816 (+0.0%)  524 (+0.0%)

  default before: 37324           2280          636
  default after:  37320 (-0.0%)   2280 (+0.0%)  636 (+0.0%)

It's not apparent in ctx because lfs3_info.name dominates (guh), but
this does save some RAM in lfs3_file_t:

  rdonly              ctx
  lfs3_file_t before: 136
  lfs3_file_t after:  112 (-17.6%)

It does add some stack cost to lfs3_file_fetch, but because this isn't
on the stack hot-path in either build, we don't really care:

  default                 code          stack          ctx
  lfs3_file_fetch before:  372            416            0
  lfs3_file_fetch after:   368 (-1.1%)    440 (+5.8%)    0 (+0.0%)
2025-06-05 18:12:51 -05:00
Christopher Haster 9eaab640e6 rdonly: Fixed lfs3_m_isrdonly shortcut leaving traversals dangling
We were relying on the previous LFS3_TSTATE_OMDIRS logic implicitly
leaving t->ot NULL when it reaches the end of the linked-list. With the
lfs3_m_isrdonly shortcut we now need to do this explicitly.

Found by test_mount_flags

Adds a bit of code to both the default and rdonly builds, but a correct
filesystem is usually preferred over a small one:

                   code          stack          ctx
  rdonly before:  10676            840          524
  rdonly after:   10680 (+0.0%)    840 (+0.0%)  524 (+0.0%)

  default before: 37320           2280          636
  default after:  37324 (+0.0%)   2280 (+0.0%)  636 (+0.0%)
2025-06-05 16:35:27 -05:00
Christopher Haster c7923ad1be rdonly: Let the compiler prune LFS3_TSTATE_OMDIRS/OBTREE
This partially reverts the LFS3_TSTATE_OMDIRS/OBTREE ifdefs, instead
adopting lfs3_m_isrdonly checks that let the compiler prune the
unreachable code paths when compiling with LFS3_RDONLY.

This adds a bit of code to both the default and rdonly builds (the
compiler isn't perfect, but simplifies the codebase:

                   code          stack          ctx
  rdonly before:  10664            840          524
  rdonly after:   10676 (+0.1%)    840 (+0.0%)  524 (+0.0%)

  default before: 37300           2280          636
  default after:  37320 (+0.1%)   2280 (+0.0%)  636 (+0.0%)

Testing the rdonly build is difficult, so minimizing the differences in
the code is quite valuable for maintenance and reliability.

As a plus, the extra ~20 bytes of code in the default build lets us
avoid traversing the omdirs when mounted LFS3_M_RDONLY. This niche
performance optimization isn't really a goal, but it's nice for
LFS3_RDONLY and LFS3_M_RDONLY to match behavior when possible.
2025-06-05 16:21:02 -05:00
Christopher Haster d791576c3f rdonly: Added missing isrdonly flag overrides when LFS3_RDONLY
I did override lfs3_o_isrdonly, but missed lfs3_m_isrdonly and
lfs3_t_isrdonly.

These aren't strictly necessary (asserts force rdonly flags to be set
correctly), but can save code by trimming unreachable code paths.

That being said, currently no observable code savings:

                  code          stack          ctx
  rdonly before: 10664            840          524
  rdonly after:  10664 (+0.0%)    840 (+0.0%)  524 (+0.0%)

But I noticed while toying around with a different way of pruning
LFS3_TSTATE_OMDIRS/OBTREE and wanted to make sure other code savings
weren't dragged in.
2025-06-05 16:21:02 -05:00
Christopher Haster e31a90d8f3 rdonly: Dropped LFS3_TSTATE_OMDIRS/OBTREE when LFS3_RDONLY
If we can't write to the filesystem, we can't out out-of-sync files, so
there's no need to traverse open file handles at all.

Saves a bit of code in LFS3_RDONLY mode:

                  code          stack          ctx
  rdonly before: 10776            840          524
  rdonly after:  10664 (-1.0%)    840 (+0.0%)  524 (+0.0%)

In theory we could also skip this check when mounted LFS3_M_RDONLY, but
checking for that flag would add code and we don't really care about
CPU-related performance here.

No code changes in default mode.
2025-06-05 16:21:02 -05:00
Christopher Haster 88eb1714b1 t: Fixed exceptional traversal errors mixing up dirty/mutated flags
This function is kinda ugly in that our failed label expects the
dirty/mutated flags to be swapped, but we only swap _after_ calling
lfs3_mtree_traverse to avoid messing up lfs3_mtree_traverse's eot logic.

Long story short, this goto failed after lfs3_mtree_traverse could end
up with drity/mutated in the wrong state.

Worst case, this can leave littlefs in a state where it thinks work was
accomplished, but only if lfs3_mtree_traverse encounters an exceptional
error (LFS3_ERR_IO? LFS3_ERR_CORRUPT?), which usually leads to emergency
actions anyways.

We probably need more testing around exceptional errors like these,
they're also the main limit to our line/branch coverage. But the work
will be tedious so for now that's a future thing.

I at least added a comment to hopefully prevent a similar regression.

Code changes minimal, humorously undoes the LFS3_RDONLY noise:

           code          stack          ctx
  before: 37304           2280          636
  after:  37300 (-0.0%)   2280 (+0.0%)  636 (+0.0%)
2025-06-05 16:21:02 -05:00
Christopher Haster 42bd130105 rdonly: Initial draft of LFS3_RDONLY
This is the new readonly flag, to be consistent with LFS3_M_RDONLY and
friends.

Note this overlaps with LFS3_YES_RDONLY in a weird way, where
LFS3_YES_RDONLY is basically just an alias for LFS3_RDONLY. For most
flags, LFS3_THING enables the _option_ of using LFS3_M_THING, with
LFS3_YES_THING implying LFS3_M_THING in all mount calls. But
LFS3_RDONLY _disables_ the option of using LFS3_M_RDWR, so it's a bit
different...

Do we really need two flags for the same thing? Not sure. But most users
probably expect LFS3_RDONLY coming from other filesystems.

Worst case this can be revisited in the planned config API rework.

---

As for the readonly code size, this is just the first draft and limited
to mostly ifdefing out all prog/write logic paths. There's some TODOs in
the code that may save a bit more (rbyd.eoff, file.b.shrub_ for
example). But the results are looking ok:

                    code           stack           ctx
  v2.11.0  rdonly:  6270             448           580
  v3-alpha rdonly: 10776 (+71.9%)    840 (+87.5%)  524 (-9.7%)

It's interesting to note most of the additional code/stack cost come
from filesystem traversal. In v2, the threaded linked-list made rdonly
traversal _incredibly_ cheap. But the extra rdwr baggage of turning
littlefs into a fully connected graph made it something to be avoided
in v3.

This hits v3 with the double whammy of:

1. Filesystem traversal is more complicated since we need to keep track
   of which btree and where in the btree we are

2. Everything needs to be tracked explicitly due to the new inverted
   state-machine driven API (no callbacks)

Note that even if we disabled the traversal APIs, lfs3_fs_usage, cksum
checking, etc, we'd still need to traverse to rebuild gstate. Otherwise
we risk showing grmed files after a powerloss.

---

This did affect the default build a little bit, due to moving things
around for nicer ifdef groupings:

                    code          stack          ctx
  default before:  37300           2280          636
  default after:   37304 (+0.0%)   2280 (+0.0%)  636 (+0.0%)
2025-06-05 16:20:41 -05:00
Christopher Haster 6eba1180c8 Big rename! Renamed lfs -> lfs3 and lfsr -> lfs3 2025-05-28 15:00:04 -05:00