7385d84df57da296b9620b52e5ad7467e098fbf9
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f51dc5c5af |
Implemented zombied file handles
A "zombie file" is a term I just made up to describe what happens when you remove a file that is currently open. To match POSIX, the opened file handle should still be available for reading/writing, even though the file doesn't really exist in the filesystem anymore. We don't have inodes, which makes this a bit more complicated, but this is where scratch files are handy again. By creating a scratch file when we remove an opened file, we preserve the mid slot for the file's sprout/shrub. We also mark the opened file as desync, so the existing orphan reclaimation circuitry kicks in when the last file handle is closed. Really the only difference between zombie files and desync files is what happens when you call lfsr_file_sync: - Desynced lfsr_file_sync => Become synced, broadcast file state. - Zombied lfsr_file_sync => Return ENOENT, you can't sync a zombie. This _is_ a bit different from POSIX, where sync on a removed file returns 0. I considered returning 0 in this case, but with all the extra behavior around sync/desync state, I figured returning ENOENT was clearer at indicating to the user sync is no longer possible. Worst case, ENOENT is not returned from sync for any other reason, so users can always treat ENOENT and 0 as the same in higher layers. The zombie file is already desynced, so close will never error. --- Implementation wise, zombies get a bit crazy. Fortunately they add little extra code, but they make up for it by adding extra subtlety. Zombie files introduce a ton of corner cases, now even directories can have zombied shrubs. This means more tests. - Seemingly unrelated operations need to be able to remove scratch files (mkdir, rename, etc). - UNCREAT state needs to be broadcasted in seemingly unrelated operations (mkdir, rename, etc). - Zombied files need to be copied over during seemingly unrelated rename operations. - And I'm sure more corner cases I'm already forgetting. One interesting tweak that simplifies things that's worth mentioning is the change to the implicitly file mid updates on rm in lfsr_mdir_commit. For non-reg files, an rm attr causes lfsr_mdir_commit to increment the mid to the next mid in the mtree. This is the correct behavior for dirs, traversals, etc. Previously, reg files were a special case that marks the mid as -1. But by changing this to also increment the mid, as well as set the zombie flag, upper layers can broadcast zombie changes by simply creating a new file and then deleting the old file in the same commit. This seems to Just Work^TM, and avoids needing to do additional state broadcasting in upper layers, which gets tricky since we may not know exactly what the new mid is post-mdir-commit. Downside: The order matters, we need to create the new file first. This violates the normal delete-then-insert order we use elsewhere to avoid overflow issues. This isn't that bad here, since we increment by at most 1. But it is something to be wary of... Still, this is much better than any other option I can think of right now. --- Uh, ignore the test_fscratch_rename* tests for now. I somehow forgot file renaming was not yet implemented... |
||
|
|
99156b5573 |
Implemented orphans resulting from closing desynced scratch files
This is a fun corner case. What happens when you close a desynced scratch file? The obvious answer seems to be just remove the scratch file in lfsr_file_close. But then what if the file is rdonly? desynced because of an error? We really shouldn't write to disk at all when closing a desync or rdonly file. This needs to be a hard rule. So the only option is to defer the work until later somehow. Fortunately, we already have several mechanisms that lead to a very nice solution. I'm very happy with this: 1. There's nothing that says our in-device grm queue needs to always match what's on-disk (we need a separate copy for xoring anyways because of the risk of leb128 encoding differences). So if we have <=2 orphans, we can just push these onto our grm. On the next write operation, the normal grm fixing code takes over and removes the pending orphans O(1). 2. If we have >2 orphans, the best we can do is mark the filesystem as having orphans, and trigger an orphan scan on the next write operation O(nlogn). But how often do you think littlefs's use cases will end up with >2 orphans? Note we also need to scan the opened-file list to make sure we're the _last_ reference to the scratch file. Otherwise we corrupt other opened file handles! --- This commit also includes a fix for a bug where the traversal mdir fell out of sync when dropping mdirs as a part of scratch file cleanup. Found when adding more tests, this would cause scratch files to go unreclaimed. |
||
|
|
0510c4b185 |
Added tests over shared scratch files
One downside of scratch files is that there are a lot of corner cases to consider. |
||
|
|
ba505c2a37 |
Implemented scratch file basics
"Scratch files" are a new file type added to solve the zero-sized
file problem. Though they have a few other uses that may be quite
valuable.
The "zero-sized file problem" is a common surprise for users, where what
seems like a simple file create+write operation:
lfs_file_open(&lfs, &file, "hi",
LFS_O_WRONLY | LFS_O_CREAT | LFS_O_EXCL);
lfs_file_write(&lfs, &file, "hello!", strlen("hello!"));
lfs_file_close(&lfs, &file);
Can end up create a zero-sized file under powerloss, breaking user
assumptions and their code.
The tricky thing is that this is actually correct behavior as defined by
POSIX. `open` with O_CREAT creats a file entry immediately, which is
initially zero-sized. And the fact that power can be lost between `open`
and `close` isn't really avoidable.
But this is a common enough footgun that it's probably worth deviating
from POSIX here.
But how to avoid zero-sized files exactly? First thought: Delay the file
creation until sync/close, tracking uncreated files in-device until
then. This solves the problem and avoids any intermediary state if we
lose power, but came with a number of headaches:
1. Since we delay file creation, we don't immediately write the filename
to disk on open. This implies we need to keep the filename allocated
in RAM until the first sync/close call.
The requirement to keep the filename allocated for new files until
first sync/close could be added to open, and with the option to call
sync immediately to save the filename (and accept the risk of
zero-sized files), I don't think it would be _that_ bad of an API.
But it would still be pretty bad. Extra bad because 1. there's no
way to warn on misuse at compile-time, 2. use-after-free bugs have a
tendency to go unnoticed annoyingly often, 3. it's a regression from
the previous API, and 4. who the heck reads the more-or-less same
`open` documentation for every filesystem they adopt.
2. Without an allocated mid, tracking files internally gets a lot
harder. The best option I could think of was to keep the opened-file
linked-list sorted by mid + (in-device) file name.
This did not feel like a great solutiona and was going to add more
code cost.
3. Handling mdir splits containing uncreated files adds another
headache. Complicated lfsr_mdir_estimate further as it needs to
decide in which mdir the uncreated files will end up, and potentially
split on a filename that isn't even created yet.
4. Since the number of uncreated files can be potentially unbounded, you
can't prevent an mdir from filling up with only uncreated files. On
disk this ends up looking like an "empty" mdir, which need specially
handling in littlefs to reclaim after powerloss.
Support for empty mdirs -- the orphaned mdir scan -- was already
added earlier. We already scan each mdir to build gstate, so it
doesn't really add much cost.
Notice that last bullet point? We already scan each mdir during mount.
Why not, instead of scanning for orphaned mdirs, scan for orphaned
files?
So this leads to the idea of "scratch files". Instead of actually
delaying file creation, fake it. Create a scratch file during open, and
on the first sync/close, convert it to a regular file. If we lose power,
scan for scratch files during mount, and remove them on first write.
Some tradeoffs:
1. The orphan scan for scratch files is a bit more expensive than for
mdirs on storage with large block sizes. We need to look at each file
entry vs just each mdir, which pushed the runtime up to O(BlogB) vs
O(B).
Though if you also consider large mtrees, the worst case is still
O(nlogn).
2. Creating intermediate scratch files adds another commit to file
creation.
This is probably not a big issue for flash, but may be more of a
concern on devices with large prog sizes.
3. Scratch files complicate unrelated mkdir/rename/etc code a bit, since
we need to consider what happens when the dest is a scratch file.
But the end result is simple. And simple is good. Both for
implementation headaches, and code size. Even if the on-disk state is
conceptually more complicated.
You may have noticed these scratch files are basically isomorphic to
just setting an "uncreated" flag on the file, and that's true. There may
have been a simpler route to end up with the design, but hey, as long as
it works.
As a plus, scratch files present a solution for a couple other things:
1. Removing an open file can become a scratch file until closed.
2. Scratch files can be used as temporary files. Open a file with
O_DESYNC and never call sync and you have yourself a temporary file.
Maybe in the future we should add O_TMPFILE to avoid the need for
unique filenames, but that is low priority.
|