Certification of Real Images through Calibrated Content Authentication

Sarim Hashmi Abdelrahman Elsayed Mohammed Talha Alam Samuele Poppi Nils Lukas

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)

Traditional detection versus calibrated resynthesis
Traditional binary detectors guess “real or fake”. Calibrated resynthesis instead asks whether any known generator can faithfully reconstruct the image, and certifies only when none can.

Abstract

A generator can reproduce authentic content, including through memorization. When authentic and generated images can be identical, image content alone cannot uniquely establish provenance. We therefore ask a narrower question: can a known generator faithfully reconstruct the query image?

If any tested generator can, synthetic provenance is plausible: the method abstains and returns the reconstruction as evidence. If none can, the method certifies the image as authentic relative to the tested generators, with a calibrated false-certification rate. Attack-aware calibration uses a stricter threshold for the evaluated bounded-perturbation setting.

We also evaluate 20 deepfake detectors against 10 generators released between December 2022 and May 2026, including white-box adversarial attacks, and find that binary detection degrades over time and collapses under bounded perturbations.

Certification is relative to the tested generators and calibration setting. Abstention does not mean an image is fake, and certification is not an unconditional proof of provenance.

Key Findings

99.5% → 76%
Generalization declines over time

Best accuracy among twenty detectors across ten generators. The leading detector shifts from PROBE to SICA to D3.

< 2%
Bounded attacks defeat detectors

Under an ℓ∞ attack with ε = 8/255, all twenty detectors fall below 2% accuracy; D3 retains the most at 1.75%.

1% FCR
Selective certification

At a 1% target false-certification rate our method still certifies authentic content, while most baselines have near-zero authentic recall.

23 / 24
Robust to transformations

Semantic transformations staying below the safety threshold. A best-of-100 PGD attacker moves the A-index only from 0.0148 to 0.0154.

Detection accuracy declines as new generators are released
Detection accuracy declines as new generators are released.

Method

  1. Reconstruct. For each generator \(G\) in the tested set, use RF-Inversion to reconstruct the query image \(x\), producing \(\tilde{x}_G\).
  2. Score. Compare image and reconstruction with pixel fidelity (PSNR), structure (SSIM), perceptual similarity (\(1-\mathrm{LPIPS}\)) and semantics (CLIP), combined into the Authenticity Index (A-index).
  3. Calibrate and certify. Set each generator’s safety threshold to the \((1-\alpha)\)-quantile of A-index scores on held-out generated samples (\(\alpha = 0.01\)). Certify only when the A-index meets the threshold for every tested generator; otherwise abstain.
$$ s(x, \tilde{x}_G) = \alpha_1\,\mathrm{PSNR} + \alpha_2\,\mathrm{SSIM} + \alpha_3\,(1 - \mathrm{LPIPS}) + \alpha_4\,\mathrm{CLIP} $$ $$ A(x, \tilde{x}_G) = \frac{\exp(-\sigma s)}{1 + \exp(-\sigma s)} $$

with \(\alpha_1 = -0.0181\), \(\alpha_2 = 1.380\), \(\alpha_3 = -4.058\), \(\alpha_4 = 8.066\) and \(\sigma = 0.9\). Faithful reconstructions receive low scores, making authenticity plausibly deniable. The security threshold applies the same quantile after attacking the calibration samples with ℓ∞-bounded PGD.

Computing the authenticity score
Computing the Authenticity Index from an image and its reconstruction.
A-index separates real and generated images
For SD3 Medium, reconstruction fidelity separates authentic and generated images.

Safety thresholds

Generator configurationThreshold
SD2.10.015
SD3 Medium0.0365
SD3.5 Medium0.0365
FLUX.1 Dev0.035
FLUX.1 Dev + Realism LoRA0.038

Certification takes 11.67 s per image per generator on one NVIDIA RTX 5000 Ada.

Results

Adversarial robustness

White-box PGD with ε = 8/255 on 2,000 images (1,000 fake, 1,000 real, 512×512). Only samples classified correctly before the attack are perturbed.

PGD collapses binary detectors but preserves separated A-index distributions
PGD perturbations collapse binary detectors, while A-index distributions remain separated.
DetectorAcc. before (%)Acc. after (%)Attack success, fake (%)Attack success, real (%)
OmniAID93.250.00100.0100.0
D383.901.7596.798.8
SICA82.000.00100.0100.0
AEROBLADE†77.801.50100.096.1
DDA77.550.20100.099.6
PROBE76.450.05100.099.9
DEAR70.500.00100.0100.0
WaRPAD†68.900.7597.8100.0

Top eight of twenty detectors by pre-attack accuracy. † marks threshold-degenerate detectors. The full table is in the repository.

Social-media study

We apply five generator-specific thresholds to the same SD3 inversion scores of 3,000 unverified Reddit images: 1,116 exceed the SD2.1 threshold, against 55–79 for the four newer configurations.

These images are unverified; threshold exceedance does not establish their true provenance.

Reddit images exceeding each generator's threshold

Preliminary video evaluation

On 100 videos from Deepfake-Eval-2024 (50 authentic, 50 generated), video detectors transfer poorly, while frame-aggregated A-index scores preserve the image-level ordering.

ModelAUCPrec.Rec.F1
GenConViT0.6150.590.490.53
FTCN0.4830.500.640.40
StyleFlow0.5090.530.420.47

Scope

BibTeX

@article{hashmi2026certification,
  title         = {Certification of Real Images through Calibrated Content Authentication},
  author        = {Hashmi, Sarim and Elsayed, Abdelrahman and Alam, Mohammed Talha and Poppi, Samuele and Lukas, Nils},
  journal       = {arXiv preprint arXiv:2610.05870},
  year          = {2026},
  eprint        = {2610.05870},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2610.05870}
}