Explore an isolated website
The Agent receives a live browser surface and a task contract. It does not receive the reference implementation, source tree, database, or private fixtures.
Evaluation protocol
The benchmark measures whether an Agent can infer and rebuild a stateful website from a bounded offline server. Dataset acceptance proves the input is trustworthy; only a later Agent run produces a benchmark score.
The Agent receives a live browser surface and a task contract. It does not receive the reference implementation, source tree, database, or private fixtures.
The Agent navigates, tests hypotheses, writes code, and produces a separately runnable clone under the same isolation rules.
Automated and multimodal judges compare observable behavior and appearance against the standard offline reference.
Scoring evidence
Matched screenshots across routes, states, and viewports; differences can be inspected with split, overlay, blink, and heatmaps.
Clicks, typing, menus, option dependencies, mutations, validation, and before/after states behave like the reference.
Success, failure, and recovery journeys reach the correct durable terminal states instead of only looking plausible.
Hidden seeds, invalid inputs, retries, persistence, reset behavior, and network isolation test behavior beyond the happy path.
Elapsed time, tool use, tokens, retries, and runtime resources contextualize the quality of the reconstruction.
Before experiments
Original website evidence is compared with the standard offline reference to certify that the benchmark input is representative, runnable, stateful, and isolated.
During experiments
The accepted offline reference is compared with a newly generated clone. This evidence produces per-dimension results and the official Agent score.
No experiment has declared a model, provider, harness, or reasoning configuration.
No Agent-produced runnable server has been submitted for comparison.
Visual, interaction, journey, robustness, efficiency, and total values are intentionally empty.
Agent traces, generated-clone screenshots, failure records, and usage metrics will appear only with real reports.