275ca0e0ec
Whoops, looks like cumulative results were overlooked when multiple bench measurements per bench were added. We were just adding all cumulative results together! This led to some very confusing bench results. The solution here is to keep track of per-measurement cumulative results via a Python dict. Which adds some memory usage, but definitely not enough to be noticeable in the context of the bench-runner.