Noise Out, Bias In

Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Sarim Hashmi* Mukul Ranjan* Abdelrahman Elsayed Muhammad Umer Sheikh Fahad Shamshad Nils Lukas

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)

* Equal contribution

⚠ Content warning: the paper and this page contain examples of stereotyped and stigmatizing content about demographic groups.
Overview of closed-loop bias injection
Closed-loop bias injection: at each denoising step a PI controller reads the target-answer probability and sets the steering strength for the next step. Weights, prompts and sampler stay unchanged.

Abstract

Masked diffusion language models (dLLMs) re-predict each masked position at every denoising step before committing a token. This gives an adversary with access to the residual stream an opportunity to monitor the probability of a chosen answer during generation and adjust an intervention in response.

We study targeted bias injection through closed-loop activation steering. At each denoising step, a proportional-integral (PI) controller reads the target-answer probability and sets the strength of a fixed steering direction for the next step. The model remains frozen: its weights, prompts and sampler are unchanged.

Highlights

1.8 → 16.7 pp
Targeted bias on BBQ

Black-target gap on ambiguous questions for LLaDA-8B-Instruct, more than three times the strongest fixed-strength baseline.

17.6% → 58.1%
SocialStigmaQA

Rate of selecting stigmatizing answers under Decode PI.

+37.2 pp
Other targets

Gap increase for Old (+22.5 pp for Arab) with no training, about 40 minutes on one GPU per target.

Feedback matters
Beyond constant steering

Constant steering matched to the controller’s mean gives a much smaller gap and ~3× the invalid-output rate.

Method

  1. Fit a steering direction offline. Using held-out contrast pairs, compute the mean residual difference at block 14 between answers naming the target group and answers naming the comparator, and normalize it to one unit direction.
  2. Control steering strength during decoding. At each denoising step, compare the target-answer probability with a setpoint. The PI controller uses this error to choose the strength αt for the next step, adding the fixed direction at every block and position.

Once the answer token commits, later controller commands no longer change it.

Direction fitting and decode-time PI control
Direction fitting and decode-time PI control.
Feedback assigns each example its own steering strength
The controller adapts to each example: it pushes harder when the target answer remains unlikely and eases off as its probability approaches the setpoint.

Results

Tgt., Cmp. and Abst. are target, comparator and abstention rates; Inv. is the invalid-output rate (kept in every denominator). Gap = target − comparator rate; Δg = change in gap vs. the unsteered base, in percentage points.

Black-target BBQ

LLaDA-8B-Instruct; 400 items, three answer rotations. Selected rows.

MethodTgt. (%)Cmp. (%)Abst. (%)Inv. (%) ↓Gap (pp) ↑Δg (pp)
Unsteered base12.811.076.30.01.8—
CAA13.710.376.00.03.3+1.6
ActAdd14.110.075.90.04.1+2.3
Mean-AcT21.918.459.40.33.5+1.7
ITI-C18.816.764.60.02.1+0.3
Open loop, tuned (α = 4)34.830.332.62.34.6+2.8
Open loop, matched to PI mean (α = 3.28)22.218.738.320.83.5+1.7
Static α per item, hindsight35.226.824.513.48.4+6.7
Layer-wise PI23.320.256.60.03.1+1.3
Decode PID (ours)29.213.248.98.816.0+14.2
Decode PI (ours)30.413.848.17.816.7+14.9

SocialStigmaQA

LLaDA-8B-Instruct; 518 items, three answer rotations. The unsteered strict-parser gap is −44.3 pp, so an all-invalid row receives Δg = +44.3 pp by construction; read gap changes alongside invalid-output rates.

MethodTgt. (%)Cmp. (%)Inv. (%) ↓Δg (pp) ↑Δgsem (pp) ↑
Unsteered base17.661.90.0——
CAA21.056.80.0+8.6+8.6
ITI-C42.542.90.0+44.0+44.0
Open loop, tuned (α = 4)52.434.13.9+62.6+63.8
Layer-wise PI35.133.50.0+45.9+45.9
Decode PID (ours)55.118.020.1+81.5+100.6
Decode PI (ours)58.118.617.4+83.8+99.9

Other targets and models

SettingMethodInv. (%) ↓Gap (pp) ↑Δg (pp)
Arab, LLaDA-8BOpen loop (α = 4)0.012.8+12.9
Decode PI (ours)9.722.3+22.5
Old, LLaDA-8BOpen loop (α = 4)19.025.8+31.3
Decode PI (ours)9.331.8+37.2
Black, Dream-7BActAdd0.06.7+3.3
Decode PI (ours)0.07.5+4.2
Black, LLaDA-MoECAA (= matched open loop)0.025.6+22.6
Decode PI (ours)0.023.3+20.3

Decode PI produces the largest gap for Arab, Old and Dream-7B; on LLaDA-MoE, CAA achieves a larger gap. Full tables are in the repository.

BibTeX

@misc{hashmi2026noiseout,
  title         = {Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering},
  author        = {Hashmi, Sarim and Ranjan, Mukul and Elsayed, Abdelrahman and Sheikh, Muhammad Umer and Shamshad, Fahad and Lukas, Nils},
  year          = {2026},
  eprint        = {2610.05894},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2610.05894}
}