The increasing realism of AI-generated images has raised critical concerns about media authenticity and the reliability of automated detection systems. Although numerous detectors have been proposed, limited effort has been devoted to analysing the cues they exploit and the factors influencing their generalization behaviour. This paper presents a preliminary explainability-based analysis of three representative AI-generated image detectors, covering frequency-domain, embedding-based, and foundation-model-based approaches. Experiments are conducted on images generated by Generative Adversarial Networks (GANs) and diffusion models across multiple semantic categories. Post-hoc explainability techniques are employed to investigate the spatial and spectral regions driving detector decisions. The results reveal distinct and complementary behaviours among the considered methods, highlighting differences in sensitivity to semantic content, frequency patterns, and generation artifacts. In addition, the impact of lightweight dissemination transformations, such as image compression, is analysed through their effects on detector explanations. Rather than introducing a new detection model, this work aims to improve the understanding of existing AI-generated image detectors by framing explainability as a tool for forensic analysis and robustness assessment.
Explainability-Driven Analysis of AI-Generated Image Detectors: Insights into Generalization and Failure Modes
Pero, Chiara;
2026-01-01
Abstract
The increasing realism of AI-generated images has raised critical concerns about media authenticity and the reliability of automated detection systems. Although numerous detectors have been proposed, limited effort has been devoted to analysing the cues they exploit and the factors influencing their generalization behaviour. This paper presents a preliminary explainability-based analysis of three representative AI-generated image detectors, covering frequency-domain, embedding-based, and foundation-model-based approaches. Experiments are conducted on images generated by Generative Adversarial Networks (GANs) and diffusion models across multiple semantic categories. Post-hoc explainability techniques are employed to investigate the spatial and spectral regions driving detector decisions. The results reveal distinct and complementary behaviours among the considered methods, highlighting differences in sensitivity to semantic content, frequency patterns, and generation artifacts. In addition, the impact of lightweight dissemination transformations, such as image compression, is analysed through their effects on detector explanations. Rather than introducing a new detection model, this work aims to improve the understanding of existing AI-generated image detectors by framing explainability as a tool for forensic analysis and robustness assessment.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


