Compare meaningful measures
Elapsed time is the interval between a run’s start and end. Activity time is the duration attributed to captured work. Human wait time is another useful dimension. Because work can overlap, these values do not always add up the way a simple progress bar suggests.
The sample run is sequential: 70 seconds of activity and 45 seconds of human wait within 115 seconds elapsed. That is an example, not a performance benchmark.
Review patterns without losing provenance
Planned views include tools and services used, observed domains, handoffs, failure hotspots, recovery outcomes, and human decisions. Every aggregation should have a clear population and a route back to its source records.
Keep absent values out of false averages
An unreported token count is not zero tokens. A cost estimate is not a settled charge. UpTrain’s proposed analytics keeps unknown values visible and states which records contribute to a calculation.
Separate assistant sessions
Reported assistant checkpoints are a different evidence class from instrumented runtime events. Comparing their activity categories can be useful, but combining them into a single verified-execution count would mislead the reviewer.