Pushed ftree tracking down into lfsr_ftree_carve
This is an attempt to reduce the overhead of on-stack snapshots during
writes, by relaxing the gaurantees provided by lfsr_file_write during
errors.
Before, thanks to the on-stack snapshot, file writes could revert to the
previous state if an error occurred. Now, on-stack snapshots are limited
to lfsr_ftree_carve, so only the state change in lfsr_ftree_carve is
reverted.
This should behave relatively predictably, since lfsr_ftree_flush calls
lfsr_ftree_carve in a normal order. If an error occurs, some, none, or
all of the data is actually written. For truncate/fruncate, there is
only one call to carve, so these remain atomic.
It's tempting to want to push this lower. If you push the atomic
operations down to the btree/bshrub level, individual btree/bshrub commits
are already atomic, so tracking on-stack bshrubs could be dropped
completely. But failures at the sub-carve level get weird! Thanks to
block crystallization and the use of order-statistic operations, the
resulting file can be quite unpredictable. For example:
on-disk: abcdefghijklmnopqrstuvwxyz
write: JKLMN
error!
on-disk: abcdstuvwxyz
With atomic carves, worst case is something like this:
on-disk: abcdefghijklmnopqrstuvwxyz
write: JKLMN
error!
on-disk: abcdefghiJKlmnopqrstuvwxyz
Thanks to file-level snapshotting, the original file can still be
recovered (and is on-disk until sync is called), but the
unpredictability is concerning.
Alternatively, if we had btree range removals, lfsr_ftree_carve could be
entirely atomic (and more efficient).
Unfortunately, btree range removals still seem like they would be
difficult to implement, and will probably be out of scope for some time.
The added code cost will also likely outweigh any savings from dropping
on-stack bshrub tracking. Still, this is probably worth looking into in
the future.
Code changes:
code stack
before: 33356 3072
after: 33006 (-1.0%) 2992 (-2.6%)
This commit is contained in:
@@ -831,6 +831,10 @@ lfs_ssize_t lfsr_file_read(lfs_t *lfs, lfsr_file_t *file,
|
||||
// Takes a buffer and size indicating the data to write. The file will not
|
||||
// actually be updated on the storage until either sync or close is called.
|
||||
//
|
||||
// If an error occurs during a write (including LFS_ERR_NOSPC), the file
|
||||
// will be marked as desynchronized, the position will be left unchanged,
|
||||
// and some, all, or none of the data may be written.
|
||||
//
|
||||
// Returns the number of bytes written, or a negative error code on failure.
|
||||
lfs_ssize_t lfs_file_write(lfs_t *lfs, lfs_file_t *file,
|
||||
const void *buffer, lfs_size_t size);
|
||||
|
||||
Reference in New Issue
Block a user