Experiment output

Agent reconstruction scores

Every score will connect back to a runnable generated clone, screenshot differences, interaction replay, terminal journey checks, and resource use. The dataset is ready; experiments have not started.

0 completed runs
Visual fidelity—
Screenshot pairs · diff maps
Interaction fidelity—
Actions · states · replay
Journey completion—
Success · failure · recovery
Robustness—
Hidden cases · isolation
Efficiency—
Time · tokens · resources

Website × agent matrix

One dataset row, no experiment columns yet

Results remain intentionally blank
Offline referenceAgent / modelGenerated cloneVisualInteractionJourneysRobustnessEfficiencyTotal
Amazon ShoppingShopping Commerce · dataset ready
Experiment not startedThe first real Agent report will create this row.
01

See the visual difference

Side-by-side, drag split, overlay, blink, and heatmap views preserve the actual screenshot pair behind visual fidelity.

ReferenceGenerated clone
02

Replay the interaction

Actions, routes, before/after states, expected outcomes, and failure points explain interaction and journey scores.

03

Audit the run

Network isolation, runtime health, time, tokens, retries, and hard failures remain visible beside the total score.

Runtime
—
Tokens
—
Failures
—