Abstract
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it.
We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6% relative to the base agent. We release our code, the benchmark, and all checkpoint results.
Method
AdvSim2Real trains three policies against one another inside WebWorld-14B, a frozen web world model that predicts the next page for any goal, page and action, including a page an adversary asks it to alter. Every trajectory is graded by a frozen LLM judge (Qwen3.8-27B). Training runs in two stages of three rounds each.
Stage 1 · curriculum and executor
- The curriculum proposes tasks: a goal and a start page.
- The executor runs each task several times; the judge grades every run.
- The curriculum is rewarded for tasks solved about half the time.
- Fresh tasks from the updated curriculum train the executor.
Stage 2 · adversary and executor
- Save several clean runs on tasks the executor already solves.
- The adversary proposes one injection and the step to insert it.
- Each clean run is replayed to that step with the injection; only success flips are rewarded.
- The executor trains on new and earlier attacks plus clean tasks.
Example
Results
Judged task completion in WebWorld-14B (%)
| Training | Executor | Clean | Adv v1 | Adv v2 | Adv v3 | Mean |
|---|---|---|---|---|---|---|
| Initial | Base | 74.89 ± 1.39 | 51.11 ± 2.34 | 48.00 ± 4.16 | 45.11 ± 1.54 | 48.07 ± 0.56 |
| Stage 1 | Capability iter 1 | 78.00 ± 1.76 | 57.11 ± 2.69 | 54.00 ± 2.91 | 53.11 ± 6.41 | 54.74 ± 3.11 |
| Capability iter 2 | 77.11 ± 0.38 | 58.89 ± 1.39 | 53.11 ± 4.54 | 54.22 ± 1.92 | 55.41 ± 1.68 | |
| Capability iter 3 | 79.33 ± 1.76 | 57.11 ± 1.39 | 52.35 ± 4.12 | 53.56 ± 3.67 | 54.34 ± 2.93 | |
| Stage 1 + 2 | Robust iter 1 | 77.33 ± 0.67 | 55.11 ± 2.14 | 52.67 ± 4.67 | 53.33 ± 2.40 | 53.70 ± 1.89 |
| Robust iter 2 | 78.00 ± 1.15 | 62.00 ± 2.00 | 54.44 ± 2.69 | 52.89 ± 1.39 | 56.44 ± 1.15 | |
| Robust iter 3 | 81.33 ± 2.31 | 62.89 ± 4.73 | 54.89 ± 3.91 | 54.67 ± 3.53 | 57.48 ± 0.56 | |
| Hosted 9B | Qwen3.5-9B | 78.22 ± 2.14 | 58.89 ± 3.36 | 57.10 ± 2.68 | 54.89 ± 1.68 | 56.96 ± 0.13 |
Mean and sample standard deviation over three rollout seeds on the 150 benchmark tasks; Mean weights Adv v1–v3 equally.
Completion under the unseen Kimi-K3 adversary (%)
| Training | Executor | Clean | Kimi-K3 |
|---|---|---|---|
| Initial | Base | 74.89 ± 1.39 | 23.00 ± 3.30 |
| Stage 1 + 2 | Robust iter 1 | 77.33 ± 0.67 | 29.67 ± 4.24 |
| Robust iter 2 | 78.00 ± 1.15 | 30.33 ± 0.47 | |
| Robust iter 3 | 81.33 ± 2.31 | 30.72 ± 3.04 |
Kimi-K3 replaces the learned adversary with the same prompt, observation and one-injection budget; two rollout seeds.
Sim-to-real transfer in a real Chromium browser (%)
| Training | Executor | Strict success | Correct fields |
|---|---|---|---|
| Initial | Base | 25.56 ± 1.02 | 52.50 ± 3.14 |
| Stage 1 | Capability iter 1 | 31.78 ± 8.34 | 61.89 ± 7.45 |
| Capability iter 2 | 43.56 ± 5.18 | 71.45 ± 4.97 | |
| Capability iter 3 | 44.44 ± 5.00 | 73.24 ± 2.24 |
Clean runs with live DOM observations and no world-model call. Strict success is a deterministic check of the submitted form; Correct fields counts the 746 target values per seed.
BibTeX
@article{hashmi2026advsim2real,
title = {AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model},
author = {Hashmi, Sarim and Ranjan, Mukul and Mishra, Kshitij and Kuznetsov, Mikhail and
Vepakomma, Praneeth and Lukas, Nils},
journal = {arXiv preprint arXiv:2610.08773},
year = {2026},
url = {https://arxiv.org/abs/2610.08773}
}