AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Sarim Hashmi1 Mukul Ranjan1 Kshitij Mishra1 Mikhail Kuznetsov2 Praneeth Vepakomma1,3 Nils Lukas1

1Mohamed bin Zayed University of Artificial Intelligence 2Amazon 3Massachusetts Institute of Technology

Overview: curriculum and adversary co-evolve with the web agent inside a frozen world model; completion over training; robustness gain over the base agent.
(a) A curriculum proposes tasks (Stage 1) and an adversary inserts one timed injection (Stage 2) into pages that a frozen world model simulates; a frozen LLM judge rewards tasks solved about half the time and success flips. (b) Judged completion on the 150 tasks over training, clean and under the learned adversaries. (c) Base agent against the final checkpoint.
74.89 → 81.33clean completion (%)
48.07 → 57.48completion under the three learned adversaries (%)
+33.6%relative gain against the unseen Kimi-K3 adversary
25.56 → 44.44strict success in a real Chromium browser (%)

Abstract

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it.

We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6% relative to the base agent. We release our code, the benchmark, and all checkpoint results.

Method

AdvSim2Real trains three policies against one another inside WebWorld-14B, a frozen web world model that predicts the next page for any goal, page and action, including a page an adversary asks it to alter. Every trajectory is graded by a frozen LLM judge (Qwen3.8-27B). Training runs in two stages of three rounds each.

Two-stage training: Stage 1 trains the curriculum and the executor; Stage 2 trains the adversary and the executor with the curriculum frozen.

Stage 1 · curriculum and executor

  1. The curriculum proposes tasks: a goal and a start page.
  2. The executor runs each task several times; the judge grades every run.
  3. The curriculum is rewarded for tasks solved about half the time.
  4. Fresh tasks from the updated curriculum train the executor.

Stage 2 · adversary and executor

  1. Save several clean runs on tasks the executor already solves.
  2. The adversary proposes one injection and the step to insert it.
  3. Each clean run is replayed to that step with the injection; only success flips are rewarded.
  4. The executor trains on new and earlier attacks plus clean tasks.

Example

The same injected notice makes the base agent click the forbidden Reset all button, while the trained agent ignores it and completes the task.
The same injected notice diverts the base agent but not the trained one (task 34, Adv v2, seed 1). The base agent clicks the forbidden Reset all right after the notice appears and is judged bad; Robust iter 3 sees the notice before five actions, never clicks it, and is judged good.

Results

Judged task completion in WebWorld-14B (%)

TrainingExecutorCleanAdv v1Adv v2Adv v3Mean
InitialBase74.89 ± 1.3951.11 ± 2.3448.00 ± 4.1645.11 ± 1.5448.07 ± 0.56
Stage 1Capability iter 178.00 ± 1.7657.11 ± 2.6954.00 ± 2.9153.11 ± 6.4154.74 ± 3.11
Capability iter 277.11 ± 0.3858.89 ± 1.3953.11 ± 4.5454.22 ± 1.9255.41 ± 1.68
Capability iter 379.33 ± 1.7657.11 ± 1.3952.35 ± 4.1253.56 ± 3.6754.34 ± 2.93
Stage 1 + 2Robust iter 177.33 ± 0.6755.11 ± 2.1452.67 ± 4.6753.33 ± 2.4053.70 ± 1.89
Robust iter 278.00 ± 1.1562.00 ± 2.0054.44 ± 2.6952.89 ± 1.3956.44 ± 1.15
Robust iter 381.33 ± 2.3162.89 ± 4.7354.89 ± 3.9154.67 ± 3.5357.48 ± 0.56
Hosted 9BQwen3.5-9B78.22 ± 2.1458.89 ± 3.3657.10 ± 2.6854.89 ± 1.6856.96 ± 0.13

Mean and sample standard deviation over three rollout seeds on the 150 benchmark tasks; Mean weights Adv v1–v3 equally.

Completion under the unseen Kimi-K3 adversary (%)

TrainingExecutorCleanKimi-K3
InitialBase74.89 ± 1.3923.00 ± 3.30
Stage 1 + 2Robust iter 177.33 ± 0.6729.67 ± 4.24
Robust iter 278.00 ± 1.1530.33 ± 0.47
Robust iter 381.33 ± 2.3130.72 ± 3.04

Kimi-K3 replaces the learned adversary with the same prompt, observation and one-injection budget; two rollout seeds.

Sim-to-real transfer in a real Chromium browser (%)

TrainingExecutorStrict successCorrect fields
InitialBase25.56 ± 1.0252.50 ± 3.14
Stage 1Capability iter 131.78 ± 8.3461.89 ± 7.45
Capability iter 243.56 ± 5.1871.45 ± 4.97
Capability iter 344.44 ± 5.0073.24 ± 2.24

Clean runs with live DOM observations and no world-model call. Strict success is a deterministic check of the submitted form; Correct fields counts the 746 target values per seed.

Analysis: Stage-1 removal ablation, Kimi-K3 against the learned adversaries, world-model versus browser outcomes, and attacked completion by skill stratum.
(a) Removing Stage 1 costs clean completion but little attacked completion. (b) Kimi-K3 requests an injection far more often than the learned adversaries. (c) World-model verdicts against browser outcomes. (d) Attacked completion by skill stratum.

BibTeX

@article{hashmi2026advsim2real,
  title   = {AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model},
  author  = {Hashmi, Sarim and Ranjan, Mukul and Mishra, Kshitij and Kuznetsov, Mikhail and
             Vepakomma, Praneeth and Lukas, Nils},
  journal = {arXiv preprint arXiv:2610.08773},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.08773}
}