scripts: Reworked dbglfs.py, adopted Lfs, Config, Gstate, etc

I'm starting to regret these reworks. They've been a big time sink. But
at least these should be much easier to extend with the future planned
auxiliary trees?

New classes:

- Bptr - A representation of littlefs's data-only block pointers.

  Extra fun is the lazily checked Bptr.__bool__ method, which should
  prevent slowing down scripts that don't actually verify checksums.

- Config - The set of littlefs config entries.

- Gstate - The set of littlefs gstate.

  I may have had too much fun with Config and Gstate. Not only do these
  provide lookup functions for config/gstate, but known config/gstate
  get lazily parsed classes that can provide easy access to the relevant
  metadata.

  These even abuse Python's __subclasses__, so all you need to do to add
  a new known config/gstate is extend the relevant Config.Config/
  Gstate.Gstate class.

  The __subclasses__ API is a weird but powerful one.

- Lfs - The big one, a high-level abstraction of littlefs itself.

  Contains subclasses for known files: Lfs.Reg, Lfs.Dir, Lfs.Stickynote,
  etc, which can be accessed by path, did+name, mid, etc. It even
  supports iterating over orphaned files, though it's expensive (but
  incredibly valuable for debugging!).

  Note that all file types can currently have attached bshrubs/btrees.
  In the existing implementation only reg files should actually end up
  with bshrubs/btrees, but the whole point of these scripts is to debug
  things that _shouldn't_ happen.

  I intentionally gave up on providing depth bounds in Lfs. Too
  complicated for something so high-level.

On noteworthy change is not recursing into directories by default. This
hopefully avoids overloading new users and matches the behavior of most
other Linux/Unix tools.

This adopts -r/--recurse/--file-depth for controlling how far to recurse
down directories, and -z/--depth/--tree-depth for controlling how far to
recurse down tree structures (mostly files). I like this API. It's
consistent with -z/--depth in the other dbg scripts, and -r/--recurse is
probably intuitive for most Linux/Unix users.

To make this work we did need to change -r/--raw -> -x/--raw. But --raw
is already a bit of a weird name for what really means "include a hex
dump".

Note that -z/--depth/--tree-depth does _not_ imply --files. Right now
only files can contain tree structures, but this will change when we get
around to adding the auxiliary trees.

This also adds the ability to specify a file path to use as the root
directory, though we need the leading slash to disambiguate file paths
and mroot addresses.

---

Also tagrepr has been tweaked to include the global/delta names,
toggleable with the optional global_ kwarg.

Rattr now has its own lazy parsers for did + name. A more organized
codebase would probably have a separate Name type, but it just wasn't
worth the hassle.

And the abstraction classes have all been tweaked to require the
explicit Rbyd.repr() function for a CLI-friendly representation. Relying
on __str__ hurt readability and debugging, especially since Python
prefers __str__ over __repr__ when printing things.
This commit is contained in:
Christopher Haster
2025-03-30 14:51:02 -05:00
parent cc20610488
commit 97b6489883
5 changed files with 5256 additions and 2264 deletions
+377 -172
View File
@@ -6,6 +6,7 @@ if __name__ == "__main__":
import bisect
import collections as co
import functools as ft
import itertools as it
import math as mt
import os
@@ -165,12 +166,17 @@ def xxd(data, width=16):
b if b >= ' ' and b <= '~' else '.'
for b in map(chr, data[i:i+width])))
def tagrepr(tag, weight=None, size=None, off=None):
# human readable tag repr
def tagrepr(tag, weight=None, size=None, *,
global_=False,
toff=None):
# null tags
if (tag & 0x6fff) == TAG_NULL:
return '%snull%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
' w%d' % weight if weight else '',
' %d' % size if size else '')
# config tags
elif (tag & 0x6f00) == TAG_CONFIG:
return '%s%s%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
@@ -185,13 +191,23 @@ def tagrepr(tag, weight=None, size=None, off=None):
else 'config 0x%02x' % (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# global-state delta tags
elif (tag & 0x6f00) == TAG_GDELTA:
return '%s%s%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
'grmdelta' if (tag & 0xfff) == TAG_GRMDELTA
else 'gdelta 0x%02x' % (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
if global_:
return '%s%s%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
'grm' if (tag & 0xfff) == TAG_GRMDELTA
else 'gstate 0x%02x' % (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
else:
return '%s%s%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
'grmdelta' if (tag & 0xfff) == TAG_GRMDELTA
else 'gdelta 0x%02x' % (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# name tags, includes file types
elif (tag & 0x6f00) == TAG_NAME:
return '%s%s%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
@@ -203,6 +219,7 @@ def tagrepr(tag, weight=None, size=None, off=None):
else 'name 0x%02x' % (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# structure tags
elif (tag & 0x6f00) == TAG_STRUCT:
return '%s%s%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
@@ -218,6 +235,7 @@ def tagrepr(tag, weight=None, size=None, off=None):
else 'struct 0x%02x' % (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# custom attributes
elif (tag & 0x6e00) == TAG_ATTR:
return '%s%sattr 0x%02x%s%s' % (
'shrub' if tag & TAG_SHRUB else '',
@@ -225,37 +243,49 @@ def tagrepr(tag, weight=None, size=None, off=None):
((tag & 0x100) >> 1) ^ (tag & 0xff),
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# alt pointers
elif tag & TAG_ALT:
return 'alt%s%s 0x%03x%s%s' % (
'r' if tag & TAG_R else 'b',
'gt' if tag & TAG_GT else 'le',
tag & 0x0fff,
' w%d' % weight if weight is not None else '',
' 0x%x' % (0xffffffff & (off-size))
if size and off is not None
' 0x%x' % (0xffffffff & (toff-size))
if size and toff is not None
else ' -%d' % size if size
else '')
# checksum tags
elif (tag & 0x7f00) == TAG_CKSUM:
return 'cksum%s%s%s%s' % (
'p' if not tag & 0xfe and tag & TAG_P else '',
' 0x%02x' % (tag & 0xff) if tag & 0xfe else '',
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# note tags
elif (tag & 0x7f00) == TAG_NOTE:
return 'note%s%s%s' % (
' 0x%02x' % (tag & 0xff) if tag & 0xff else '',
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# erased-state checksum tags
elif (tag & 0x7f00) == TAG_ECKSUM:
return 'ecksum%s%s%s' % (
' 0x%02x' % (tag & 0xff) if tag & 0xff else '',
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# global-checksum delta tags
elif (tag & 0x7f00) == TAG_GCKSUMDELTA:
return 'gcksumdelta%s%s%s' % (
' 0x%02x' % (tag & 0xff) if tag & 0xff else '',
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
if global_:
return 'gcksum%s%s%s' % (
' 0x%02x' % (tag & 0xff) if tag & 0xff else '',
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
else:
return 'gcksumdelta%s%s%s' % (
' 0x%02x' % (tag & 0xff) if tag & 0xff else '',
' w%d' % weight if weight else '',
' %s' % size if size is not None else '')
# unknown tags
else:
return '0x%04x%s%s' % (
tag,
@@ -302,6 +332,29 @@ class TreeBranch(co.namedtuple('TreeBranch', ['a', 'b', 'depth', 'color'])):
def __ge__(self, other):
return (self.depth, self.a, self.b) >= (other.depth, other.a, other.b)
# apply a function to a/b while trying to avoid copies
def map(self, filter_, map_=None):
if map_ is None:
filter_, map_ = None, filter_
a = self.a
if filter_ is None or filter_(a):
a = map_(a)
b = self.b
if filter_ is None or filter_(b):
b = map_(b)
if a != self.a or b != self.b:
return self.__class__(
a if a != self.a else self.a,
b if b != self.b else self.b,
self.depth,
self.color)
else:
return self
# render some nice ascii trees
def treerepr(tree, x, depth=None, color=False):
# find the max depth from the tree
if depth is None:
@@ -359,9 +412,15 @@ def pathdelta(a, b):
a = list(a)
i = 0
for a_, b_ in zip(a, b):
if type(a_) == type(b_) and a_ == b_:
i += 1
else:
try:
if type(a_) == type(b_) and a_ == b_:
i += 1
else:
break
# treat exceptions here as failure to match, most likely
# the compared types are incompatible, it's the caller's
# problem
except Exception:
break
return [(i+j, a_) for j, a_ in enumerate(a[i:])]
@@ -375,17 +434,17 @@ class Bd:
self.block_count = block_count
def __repr__(self):
return '<%s %sx%s>' % (
self.__class__.__name__,
self.block_size,
self.block_count)
return '<%s %s>' % (self.__class__.__name__, self.repr())
def repr(self):
return 'bd %sx%s' % (self.block_size, self.block_count)
def read(self, size=-1):
return self.f.read(size)
def seek(self, block, off, whence=0):
def seek(self, block, off=0, whence=0):
pos = self.f.seek(block*self.block_size + off, whence)
return pos // block_size, pos % block_size
return pos // self.block_size, pos % self.block_size
def readblock(self, block):
self.f.seek(block*self.block_size)
@@ -393,22 +452,40 @@ class Bd:
# tagged data in an rbyd
class Rattr:
def __init__(self, tag, weight, block, toff, off, data):
def __init__(self, tag, weight, blocks, toff, tdata, data):
self.tag = tag
self.weight = weight
self.block = block
if isinstance(blocks, int):
self.blocks = [blocks]
else:
self.blocks = list(blocks)
self.toff = toff
self.off = off
self.tdata = tdata
self.data = data
@property
def block(self):
return self.blocks[0]
@property
def tsize(self):
return len(self.tdata)
@property
def off(self):
return self.toff + len(self.tdata)
@property
def size(self):
return len(self.data)
def __repr__(self):
return '<%s %s>' % (self.__class__.__name__, self)
def __bytes__(self):
return self.data
def __str__(self):
def __repr__(self):
return '<%s %s>' % (self.__class__.__name__, self.repr())
def repr(self):
return tagrepr(self.tag, self.weight, self.size)
def __iter__(self):
@@ -424,14 +501,42 @@ class Rattr:
def __hash__(self):
return hash((self.tag, self.weight, self.data))
# convenience for did/name access
def _parse_name(self):
# note we return a null name for non-name tags, this is so
# vestigial names in btree nodes act as a catch-all
if (self.tag & 0xff00) != TAG_NAME:
did = 0
name = b''
else:
did, d = fromleb128(self.data)
name = self.data[d:]
# cache both
self.did = did
self.name = name
@ft.cached_property
def did(self):
self._parse_name()
return self.did
@ft.cached_property
def name(self):
self._parse_name()
return self.name
class Ralt:
def __init__(self, tag, weight, block, toff, off, jump,
def __init__(self, tag, weight, blocks, toff, tdata, jump,
color=None, followed=None):
self.tag = tag
self.weight = weight
self.block = block
if isinstance(blocks, int):
self.blocks = [blocks]
else:
self.blocks = list(blocks)
self.toff = toff
self.off = off
self.tdata = tdata
self.jump = jump
if color is not None:
@@ -440,15 +545,27 @@ class Ralt:
self.color = 'r' if tag & TAG_R else 'b'
self.followed = followed
@property
def block(self):
return self.blocks[0]
@property
def tsize(self):
return len(self.tdata)
@property
def off(self):
return self.toff + len(self.tdata)
@property
def joff(self):
return self.toff - self.jump
def __repr__(self):
return '<%s %s>' % (self.__class__.__name__, self)
return '<%s %s>' % (self.__class__.__name__, self.repr())
def __str__(self):
return tagrepr(self.tag, self.weight, self.jump, self.toff)
def repr(self):
return tagrepr(self.tag, self.weight, self.jump, toff=self.toff)
def __iter__(self):
return iter((self.tag, self.weight, self.jump))
@@ -466,19 +583,20 @@ class Ralt:
# our core rbyd type
class Rbyd:
def __init__(self, data, blocks, trunk, weight, rev, eoff, cksum, *,
def __init__(self, blocks, trunk, weight, rev, eoff, cksum, data, *,
gcksumdelta=None,
corrupt=False):
if isinstance(blocks, int):
blocks = [blocks]
self.data = data
self.blocks = list(blocks)
self.blocks = [blocks]
else:
self.blocks = list(blocks)
self.trunk = trunk
self.weight = weight
self.rev = rev
self.eoff = eoff
self.cksum = cksum
self.data = data
self.gcksumdelta = gcksumdelta
self.corrupt = corrupt
@@ -495,10 +613,10 @@ class Rbyd:
self.trunk)
def __repr__(self):
return '<%s %s w%s>' % (
self.__class__.__name__,
self.addr(),
self.weight)
return '<%s %s>' % (self.__class__.__name__, self.repr())
def repr(self):
return 'rbyd %s w%s' % (self.addr(), self.weight)
def __bool__(self):
return not self.corrupt
@@ -514,11 +632,11 @@ class Rbyd:
return hash((frozenset(self.blocks), self.trunk))
@classmethod
def fetch(cls, bd, blocks, trunk=None, cksum=None):
def fetch(cls, bd, blocks, trunk=None):
# multiple blocks? unfortunately this must be a list
if isinstance(blocks, list):
# fetch all blocks
rbyds = [cls.fetch(bd, block, trunk, cksum) for block in blocks]
rbyds = [cls.fetch(bd, block, trunk) for block in blocks]
# determine most recent revision
i = 0
for i_, rbyd in enumerate(rbyds):
@@ -534,6 +652,9 @@ class Rbyd:
rbyd.blocks += tuple(
rbyds[(i+1+j) % len(rbyds)].block
for j in range(len(rbyds)-1))
# and patch the gcksumdelta if we have one
if rbyd.gcksumdelta is not None:
rbyd.gcksumdelta.blocks = rbyd.blocks
return rbyd
block = blocks
@@ -546,9 +667,9 @@ class Rbyd:
else block[1] if isinstance(block, tuple)
else None)
# bd can be either a bd reference or preread data
# bd can be either a bd reference or a preread block
#
# preread data can be useful for avoiding race conditions
# preread blocks can be useful for avoiding race conditions
# with cksums and shrubs
if isinstance(bd, Bd):
# seek/read the block
@@ -558,9 +679,9 @@ class Rbyd:
# fetch the rbyd
rev = fromle32(data[0:4])
cksum_ = 0
cksum__ = crc32c(data[0:4])
cksum___ = cksum__
cksum = 0
cksum_ = crc32c(data[0:4])
cksum__ = cksum_
perturb = False
eoff = 0
eoff_ = None
@@ -576,10 +697,10 @@ class Rbyd:
while j_ < len(data) and (not trunk or eoff <= trunk):
# read next tag
v, tag, w, size, d = fromtag(data[j_:])
if v != parity(cksum___):
if v != parity(cksum__):
break
cksum___ ^= 0x00000080 if v else 0
cksum___ = crc32c(data[j_:j_+d], cksum___)
cksum__ ^= 0x00000080 if v else 0
cksum__ = crc32c(data[j_:j_+d], cksum__)
j_ += d
if not tag & TAG_ALT and j_ + size > len(data):
break
@@ -587,22 +708,23 @@ class Rbyd:
# take care of cksums
if not tag & TAG_ALT:
if (tag & 0xff00) != TAG_CKSUM:
cksum___ = crc32c(data[j_:j_+size], cksum___)
cksum__ = crc32c(data[j_:j_+size], cksum__)
# found a gcksumdelta?
if (tag & 0xff00) == TAG_GCKSUMDELTA:
gcksumdelta_ = Rattr(tag, w,
block, j_-d, d, data[j_:j_+size])
gcksumdelta_ = Rattr(tag, w, block, j_-d,
data[j_-d:j_],
data[j_:j_+size])
# found a cksum?
else:
# check cksum
cksum____ = fromle32(data[j_:j_+4])
if cksum___ != cksum____:
cksum___ = fromle32(data[j_:j_+4])
if cksum__ != cksum___:
break
# commit what we have
eoff = eoff_ if eoff_ else j_ + size
cksum_ = cksum__
cksum = cksum_
trunk_ = trunk__
weight = weight_
gcksumdelta = gcksumdelta_
@@ -610,7 +732,7 @@ class Rbyd:
# update perturb bit
perturb = tag & TAG_P
# revert to data cksum and perturb
cksum___ = cksum__ ^ (0xfca42daf if perturb else 0)
cksum__ = cksum_ ^ (0xfca42daf if perturb else 0)
# evaluate trunks
if (tag & 0xf000) != TAG_CKSUM:
@@ -634,7 +756,7 @@ class Rbyd:
if trunk and j_ + size > trunk:
eoff_ = j_ + size
eoff = eoff_
cksum_ = cksum___ ^ (
cksum = cksum__ ^ (
0xfca42daf if perturb else 0)
trunk_ = trunk__
weight = weight_
@@ -642,23 +764,34 @@ class Rbyd:
trunk___ = 0
# update canonical checksum, xoring out any perturb state
cksum__ = cksum___ ^ (0xfca42daf if perturb else 0)
cksum_ = cksum__ ^ (0xfca42daf if perturb else 0)
if not tag & TAG_ALT:
j_ += size
# cksum mismatch?
if cksum is not None and cksum_ != cksum:
return cls(data, block, 0, 0, rev, 0, cksum_,
corrupt=True)
return cls(data, block, trunk_, weight, rev, eoff, cksum_,
return cls(block, trunk_, weight, rev, eoff, cksum, data,
gcksumdelta=gcksumdelta,
corrupt=not trunk_)
@classmethod
def fetchck(cls, bd, blocks, trunk, weight, cksum):
# try to fetch the rbyd normally
rbyd = cls.fetch(bd, blocks, trunk)
# cksum mismatch? trunk/weight mismatch?
if (rbyd.cksum != cksum
or rbyd.trunk != trunk
or rbyd.weight != weight):
# mark as corrupt and keep track of expected trunk/weight
rbyd.corrupt = True
rbyd.trunk = trunk
rbyd.weight = weight
return rbyd
def lookupnext(self, rid, tag=None, *,
path=False):
if not self:
if not self or rid >= self.weight:
return None, None, *(([],) if path else ())
tag = max(tag or 0, 0x1)
@@ -694,7 +827,8 @@ class Rbyd:
color = 'b'
path_.append(Ralt(
alt, w, self.block, j+jump, j+jump+d, jump,
alt, w, self.blocks, j+jump,
self.data[j+jump:j+jump+d], jump,
color=color,
followed=True))
@@ -716,7 +850,8 @@ class Rbyd:
color = 'b'
path_.append(Ralt(
alt, w, self.block, j-d, j, jump,
alt, w, self.blocks, j-d,
self.data[j-d:j], jump,
color=color,
followed=False))
@@ -730,7 +865,8 @@ class Rbyd:
return None, None, *(([],) if path else ())
return (rid_,
Rattr(tag_, w_, self.block, j, j+d,
Rattr(tag_, w_, self.blocks, j,
self.data[j:j+d],
self.data[j+d:j+d+jump]),
*((path_,) if path else ()))
@@ -784,7 +920,7 @@ class Rbyd:
yield rid, name, *path_
rid += 1
def rattrs_(self, rid=None, *,
def rattrs_(self, rid=None, tag=None, mask=None, *,
path=False):
if rid is None:
rid, tag = -1, 0
@@ -798,24 +934,31 @@ class Rbyd:
yield rid, rattr, *path_
tag = rattr.tag
else:
tag = 0
if tag is None:
tag, mask = 0, 0xffff
if mask is None:
mask = 0
tag_ = max((tag & ~mask) - 1, 0)
while True:
rid_, rattr, *path_ = self.lookupnext(rid, tag+0x1,
rid_, rattr_, *path_ = self.lookupnext(rid, tag_+0x1,
path=path)
# found end of tree?
if rid_ is None or rid_ != rid:
if (rid_ is None
or rid_ != rid
or (rattr_.tag & ~mask) != (tag & ~mask)):
break
yield rattr, *path_
tag = rattr.tag
yield rattr_, *path_
tag_ = rattr_.tag
def rattrs(self, rid=None, *,
def rattrs(self, rid=None, tag=None, mask=None, *,
path=False):
if rid is None:
yield from self.rattrs_(rid,
yield from self.rattrs_(rid, tag, mask,
path=path)
else:
for rattr, *path_ in self.rattrs_(rid,
for rattr, *path_ in self.rattrs_(rid, tag, mask,
path=path):
if path:
yield rattr, *path_
@@ -828,33 +971,25 @@ class Rbyd:
# lookup by name
def namelookup(self, did, name):
# binary search
best = (False, None, None, None, None)
best = None, None
lower = 0
upper = self.weight
while lower < upper:
rid, rattr = self.lookupnext(lower + (upper-1-lower)//2)
rid, name_ = self.lookupnext(
lower + (upper-1-lower)//2)
if rid is None:
break
# treat vestigial names as a catch-all
if ((rattr.tag == TAG_NAME and rid-(rattr.weight-1) == 0)
or (rattr.tag & 0xff00) != TAG_NAME):
did_ = 0
name_ = b''
else:
did_, d = fromleb128(rattr.data)
name_ = rattr.data[d:]
# bisect search space
if (did_, name_) > (did, name):
upper = rid-(w-1)
elif (did_, name_) < (did, name):
if (name_.did, name_.name) > (did, name):
upper = rid-(name_.weight-1)
elif (name_.did, name_.name) < (did, name):
lower = rid + 1
# keep track of best match
best = (False, rid, rattr)
best = rid, name_
else:
# found a match
return True, rid, rattr
return rid, name_
return best
@@ -1006,10 +1141,10 @@ class Btree:
return self.rbyd.addr()
def __repr__(self):
return '<%s %s w%s>' % (
self.__class__.__name__,
self.addr(),
self.weight)
return '<%s %s>' % (self.__class__.__name__, self.repr())
def repr(self):
return 'btree %s w%s' % (self.addr(), self.weight)
def __eq__(self, other):
return self.rbyd == other.rbyd
@@ -1021,16 +1156,40 @@ class Btree:
return hash(self.rbyd)
@classmethod
def fetch(cls, bd, blocks, trunk=None, cksum=None):
# we need a real bd reference here
def fetch(cls, bd, blocks, trunk=None):
# bd can either be a bd reference or a tuple of bd + data to
# avoid rereads, but we need a real bd reference somehow
if isinstance(bd, tuple):
bd, data = bd
else:
bd, data = bd, bd
assert isinstance(bd, Bd)
rbyd = Rbyd.fetch(bd, blocks, trunk, cksum)
# rbyd fetch does most of the work here
rbyd = Rbyd.fetch(data, blocks, trunk)
return cls(bd, rbyd)
@classmethod
def fetchck(cls, bd, blocks, trunk, weight, cksum):
# bd can either be a bd reference or a tuple of bd + data to
# avoid rereads, but we need a real bd reference somehow
if isinstance(bd, tuple):
bd, data = bd
else:
bd, data = bd, bd
assert isinstance(bd, Bd)
# rbyd fetchck does most of the work here
rbyd = Rbyd.fetchck(data, blocks, trunk, weight, cksum)
return cls(bd, rbyd)
def lookupleaf(self, bid, *,
path=None,
depth=None):
if not self or bid >= self.weight:
return (None, None, None, None,
*(([],) if path else ()))
rbyd = self.rbyd
rid = bid
depth_ = 1
@@ -1059,11 +1218,8 @@ class Btree:
if branch_ is not None and (
not depth or depth_ < depth):
block, trunk, cksum = frombranch(branch_.data)
rbyd = Rbyd.fetch(self.bd, block, trunk, cksum)
# keep track of expected trunk/weight if corrupted
if not rbyd:
rbyd.trunk = trunk
rbyd.weight = name_.weight
rbyd = Rbyd.fetchck(self.bd, block, trunk, name_.weight,
cksum)
rid -= (rid_-(name_.weight-1))
depth_ += 1
@@ -1072,22 +1228,51 @@ class Btree:
return (bid + (rid_-rid), rbyd, rid_, name_,
*((path_,) if path else ()))
def lookup(self, bid, tag=None, mask=None, *,
# the non-leaf variants discard the rbyd info, these can be a bit
# more convenient, but at a performance cost
def lookupnext(self, bid, *,
path=None,
depth=None):
# just discard the rbyd info
bid, rbyd, rid, name, *path_ = self.lookupleaf(bid,
path=path,
depth=depth)
return bid, name, *path_
def lookup_(self, bid, tag=None, mask=None, *,
path=False,
depth=None):
# lookup rbyd in btree
#
# note this function expects bid to be known, use lookupnext
# first if you don't care about the exact bid (or better yet,
# lookupleaf and call lookup on the returned rbyd)
#
# this matches rbyd's lookup behavior, which needs a known rid
# to avoid a double lookup
bid_, rbyd_, rid_, name_, *path_ = self.lookupleaf(bid,
path=path,
depth=depth)
if bid_ is None:
return None, None, None, None, None, *path_
if bid_ is None or bid_ != bid:
return None, *path_
# lookup tag in rbyd
rattr_ = rbyd_.lookup(rid_, tag, mask)
if rattr_ is None:
return None, None, None, None, None, *path_
return None, *path_
return bid_, rbyd_, rid_, name_, rattr_, *path_
return rattr_, *path_
def lookup(self, bid, tag=None, mask=None, *,
path=False,
depth=None):
rattr, *path_ = self.lookup_(bid, tag, mask,
path=path,
depth=depth)
if path:
return rattr, *path_
else:
return rattr
def __getitem__(self, key):
if not isinstance(key, tuple):
@@ -1099,7 +1284,7 @@ class Btree:
if not isinstance(key, tuple):
key = (key,)
return self.lookup(*key)[0] is not None
return self.lookup_(*key)[0] is not None
# note leaves only iterates over leaf rbyds, whereas traverse
# traverses all rbyds
@@ -1108,7 +1293,8 @@ class Btree:
depth=None):
# include our root rbyd even if the weight is zero
if self.weight == 0:
yield -1, self.rbyd, *(([],) if path else())
yield -1, self.rbyd, *(([],) if path else ())
return
bid = 0
while True:
@@ -1139,21 +1325,27 @@ class Btree:
yield bid_, rbyd_, *((path_[:d],) if path else ())
ptrunk_ = trunk_
# note bids/rattrs do _not_ include corrupt btree nodes!
def bids(self, *,
leaves=False,
path=False,
depth=None):
for bid, rbyd, *path_ in self.leaves(
path=path,
depth=depth):
for rid, name in rbyd.rids():
yield (bid-(rbyd.weight-1) + rid,
rbyd, rid, name,
*((path_[0]+[
(bid-(rbyd.weight-1) + rid,
rbyd, rid, name)],)
if path else ()))
bid_ = bid-(rbyd.weight-1) + rid
if leaves:
yield (bid_, rbyd, rid, name,
*((path_[0]+[(bid_, rbyd, rid, name)],)
if path else ()))
else:
yield (bid_, name,
*((path_[0]+[(bid_, rbyd, rid, name)],)
if path else ()))
def rattrs_(self, bid=None, *,
def rattrs_(self, bid=None, tag=None, mask=None, *,
leaves=False,
path=False,
depth=None):
if bid is None:
@@ -1161,48 +1353,68 @@ class Btree:
path=path,
depth=depth):
for rid, name in rbyd.rids():
bid_ = bid-(rbyd.weight-1) + rid
for rattr in rbyd.rattrs(rid):
yield (bid-(rbyd.weight-1) + rid,
rbyd, rid, name, rattr,
*((path_[0]+[
(bid-(rbyd.weight-1) + rid,
rbyd, rid, name)],)
if path else ()))
if leaves:
yield (bid_, rbyd, rid, rattr,
*((path_[0]+[(bid_, rbyd, rid, name)],)
if path else ()))
else:
yield (bid_, rattr,
*((path_[0]+[(bid_, rbyd, rid, name)],)
if path else ()))
else:
bid, rbyd, rid, name, *path_ = self.lookupleaf(bid,
path=path,
depth=depth)
for rattr in rbyd.rattrs(rid):
yield rattr, *path_
if bid is None:
return
def rattrs(self, bid=None, *,
for rattr in rbyd.rattrs(rid, tag, mask):
if leaves:
yield rbyd, rid, rattr, *path_
else:
yield rattr, *path_
def rattrs(self, bid=None, tag=None, mask=None, *,
leaves=False,
path=False,
depth=None):
if bid is None:
yield from self.rattrs_(bid,
if bid is None or leaves or path:
yield from self.rattrs_(bid, tag, mask,
leaves=leaves,
path=path,
depth=depth)
else:
for rattr, *path_ in self.rattrs_(bid,
for rattr, *path_ in self.rattrs_(bid, tag, mask,
leaves=leaves,
path=path,
depth=depth):
if path:
yield rattr, *path_
else:
yield rattr
yield rattr
def __iter__(self):
return self.rattrs()
# lookup by name
def namelookup(self, did, name, *,
def namelookupleaf(self, did, name, *,
path=None,
depth=None):
rbyd = self.rbyd
bid = 0
depth_ = 1
path_ = []
while True:
found_, rid_, name_ = rbyd.namelookup(did, name)
# corrupt branch?
if not rbyd:
return (bid+(rbyd.weight-1), rbyd, rbyd.weight-1, None,
*((path_,) if path else ()))
rid_, name_ = rbyd.namelookup(did, name)
# keep track of path
if path:
path_.append((bid + rid_, rbyd, rid_, name_))
# find branch tag if there is one
branch_ = rbyd.lookup(rid_, TAG_BRANCH, 0x3)
@@ -1210,21 +1422,27 @@ class Btree:
# found another branch
if branch_ is not None and (
not depth or depth_ < depth):
block, trunk, cksum = frombranch(branch_.data)
rbyd = Rbyd.fetchck(self.bd, block, trunk, name_.weight,
cksum)
# update our bid
bid += rid_ - (name_.weight-1)
block, trunk, cksum = frombranch(branch_.data)
rbyd = Rbyd.fetch(self.bd, block, trunk, cksum)
# keep track of expected trunk/weight if corrupted
if not rbyd:
rbyd.trunk = trunk
rbyd.weight = name_.weight
depth_ += 1
# found best match
else:
return found_, bid + rid_, rbyd, rid_, name_
return (bid + rid_, rbyd, rid_, name_,
*((path_,) if path else ()))
def namelookup(self, bid, *,
path=None,
depth=None):
# just discard the rbyd info
bid, rbyd, rid, name, *path_ = self.namelookupleaf(did, name,
path=path,
depth=depth)
return bid, name, *path_
# create an rbyd tree for debugging
def _tree_rtree(self, *,
@@ -1315,22 +1533,10 @@ class Btree:
roots[bids[i]] = t
# remap branches to leaf-roots
tree_ = set()
for t in tree:
if t.a[1] == d and t.a[0] in roots:
t = TreeBranch(
roots[t.a[0]].a,
t.b,
t.depth,
t.color)
if t.b[1] == d and t.b[0] in roots:
t = TreeBranch(
t.a,
roots[t.b[0]].a,
t.depth,
t.color)
tree_.add(t)
tree = tree_
tree = {t.map(
lambda x: x[1] == d and x[0] in roots,
lambda x: roots[x[0]].a)
for t in tree}
return tree
@@ -1343,7 +1549,7 @@ class Btree:
tree = set()
root = None
branches = {}
for bid, rbyd, rid, name, path in self.bids(
for bid, name, path in self.bids(
path=True,
depth=depth):
# create branch for each jump in path
@@ -1463,7 +1669,7 @@ def main(disk, roots=None, *,
if rattr.weight > 1
else bid if rattr.weight > 0
else '',
21+w_width, rattr,
21+w_width, rattr.repr(),
next(xxd(rattr.data, 8), '')
if not args.get('raw')
and not args.get('no_truncate')
@@ -1472,8 +1678,7 @@ def main(disk, roots=None, *,
# show on-disk encoding of tags/data
if args.get('raw'):
for o, line in enumerate(xxd(
rbyd.data[rattr.toff:rattr.off])):
for o, line in enumerate(xxd(rattr.tdata)):
print('%9s: %*s%*s %s' % (
'%04x' % (rattr.toff + o*16),
t_width, '',
@@ -1553,7 +1758,7 @@ if __name__ == "__main__":
default='auto',
help="When to use terminal colors. Defaults to 'auto'.")
parser.add_argument(
'-r', '--raw',
'-x', '--raw',
action='store_true',
help="Show the raw data including tag encodings.")
parser.add_argument(
@@ -1581,7 +1786,7 @@ if __name__ == "__main__":
nargs='?',
type=lambda x: int(x, 0),
const=0,
help="Depth of tree to show.")
help="Depth of the btree to show.")
parser.add_argument(
'-e', '--error-on-corrupt',
action='store_true',