Agent website reconstruction benchmark

Can an agent rebuild a website it can only explore?

WebsiteBench gives an agent an isolated offline website—not its source code. The agent must interact with it, infer its behavior, and produce a new runnable clone. This viewer makes the input website, the generated clone, and every score inspectable.

Dataset construction · active Agent experiments · not started
offline.amazon.local
Amazon Shopping offline reference website
Amazon Shopping · worked exampleOpen evidence →
01
Available now

Offline reference website

A stateful, isolated server the agent can inspect only through interaction.

02
Reserved

Agent exploration

Browser actions, visited states, tool traces, and resource use will appear here.

03
Reserved

Generated clone

The new runnable server produced by each evaluated agent.

04
Reserved

Multimodal scoring

Screenshots, interaction replay, terminal states, robustness, and efficiency.

Corpus roadmap

More website categories will be added here

1 available · 3 reserved
01
Amazon Shopping

Shopping Commerce

Accepted
02
Future website

Category and scope not assigned

Reserved
03
Future website

Category and scope not assigned

Reserved
04
Future website

Category and scope not assigned

Reserved

Future experiment output

Every agent run will stay explainable

Open scoring workspace →
Visual fidelity
Awaiting first agent run
Interaction fidelity
Awaiting first agent run
Journey completion
Awaiting first agent run
Robustness
Awaiting first agent run
Efficiency
Awaiting first agent run
No agent reconstruction has been evaluated yet.

Model names, generated clones, replays, and scores remain blank until real experiment reports are added.