Self-blended images are widely used to train face-swap detectors, but primarily capture blending artifacts.
We investigate whether adding illumination inconsistencies improves detection.
Temporal Self-Blended Images (T-SBI) transfer lighting statistics between frames of the same video, with the mismatch controlled by luminance difference (ΔL).
Using five training regimes and a three-seed comparison of high- and low-ΔL training, we find no evidence of illumination-specific improvements.
AUC differences remain within seed variability across four datasets, and an analysis of 506,328 attribute-binned samples shows no preferential reduction in errors under harsh lighting.
Instead, T-SBI shifts prediction scores, changing optimal thresholds by approximately 0.34 on FaceForensics++ and 0.30 on Celeb-DF, making comparisons at a fixed threshold misleading.
However, T-SBI improves robustness to heavy JPEG compression on DFDC (AUC 0.780 versus 0.696), potentially reflecting greater reliance on low-frequency cues.
These findings highlight the importance of evaluating training methods against their intended targets and accounting for threshold effects.