DACH-Web-Agent-Bench: an open benchmark for the German web
What DACH-Web-Agent-Bench measures: 83 tasks across ten dimensions against a fixture site, Apache-2.0 licensed and rerunnable by anyone, including competitors.
Last updated:
What the benchmark is
DACH-Web-Agent-Bench is a benchmark for browser agents on the German-language web: 83 tasks across ten dimensions, run against a fixture site that ships with it. No task depends on a live third-party website, so a run today and a run next year measure the same thing.
The benchmark is Apache-2.0 and has zero dependencies on the rest of Browserberg — that is the point, not an accident. A competitor can check it out, wire up their own agent and rerun every task. Browserberg's own agent enters through the same public adapter interface a third party would use; there is no privileged path.
Why it exists
Generic web-agent benchmarks miss what the German web actually looks like: consent layers in front of everything, legacy portals that predate the frameworks agents are tuned on, and legal text an agent must read rather than skim. The task set is built from that reality.
Refusal is scored in the same total in both directions. An agent earns points for doing what should be done and for declining what must be declined — a benchmark that only rewards action teaches agents to click through warnings.
Published third-party results do not exist yet. Rather than quote our own numbers at you, we point at the design: the benchmark is open, self-contained and cheap to inspect, so the honest move is to run it yourself.