Experiment output

Agent reconstruction scores

Every score will connect back to a runnable generated clone, screenshot differences, interaction replay, terminal journey checks, and resource use. The dataset is ready; experiments have not started.

0 completed runs
Visual fidelity
Screenshot pairs · diff maps
Interaction fidelity
Actions · states · replay
Journey completion
Success · failure · recovery
Robustness
Hidden cases · isolation
Efficiency
Time · tokens · resources

Website × agent matrix

One dataset row, no experiment columns yet

Results remain intentionally blank
Offline referenceAgent / modelGenerated cloneVisualInteractionJourneysRobustnessEfficiencyTotal
Amazon ShoppingShopping Commerce · dataset ready
Experiment not startedThe first real Agent report will create this row.
01

See the visual difference

Side-by-side, drag split, overlay, blink, and heatmap views preserve the actual screenshot pair behind visual fidelity.

ReferenceGenerated clone
02

Replay the interaction

Actions, routes, before/after states, expected outcomes, and failure points explain interaction and journey scores.

03

Audit the run

Network isolation, runtime health, time, tokens, retries, and hard failures remain visible beside the total score.

Runtime
Tokens
Failures