Experiment output
Agent reconstruction scores
Every score will connect back to a runnable generated clone, screenshot differences, interaction replay, terminal journey checks, and resource use. The dataset is ready; experiments have not started.
Website × agent matrix
One dataset row, no experiment columns yet
| Offline reference | Agent / model | Generated clone | Visual | Interaction | Journeys | Robustness | Efficiency | Total |
|---|---|---|---|---|---|---|---|---|
| Amazon ShoppingShopping Commerce · dataset ready | Experiment not startedThe first real Agent report will create this row. | |||||||
See the visual difference
Side-by-side, drag split, overlay, blink, and heatmap views preserve the actual screenshot pair behind visual fidelity.
ReferenceGenerated clone
Replay the interaction
Actions, routes, before/after states, expected outcomes, and failure points explain interaction and journey scores.
Audit the run
Network isolation, runtime health, time, tokens, retries, and hard failures remain visible beside the total score.
- Runtime
- —
- Tokens
- —
- Failures
- —