Benchmark: measuring agents on the German web
The DACH-Web-Agent-Bench section: what the benchmark covers, how to run it against your own agent adapter, and how its scoring works.
DACH-Web-Agent-Bench is Browserberg's open benchmark for browser agents on the German-language web — 83 tasks, ten dimensions, a self-contained fixture site, and an adapter interface that treats our agent and yours identically. This section covers what it measures, how to run it from a checkout, and how the scoring is designed so you never have to take anyone's word for a result.
All pages in this section
-
1
DACH-Web-Agent-Bench: an open benchmark for the German web
What DACH-Web-Agent-Bench measures: 83 tasks across ten dimensions against a fixture site, Apache-2.0 licensed and rerunnable by anyone, including competitors.
-
2
Running DACH-Web-Agent-Bench against your own agent
How to run the benchmark from a repository checkout: the CLI commands, the agent adapter interface, and what a run against a real model costs you.
-
3
Benchmark methodology: ten dimensions and the oracle design
The ten task dimensions, why the fixture server is the only oracle a scorer trusts, and how refusal counts in the same total as action.