The increasing realism of AI-generated images has raised critical concerns about media authenticity and the reliability of automated detection systems. Although numerous detectors have been proposed, limited effort has been devoted to analysing the cues they exploit and the factors influencing their generalization behaviour. This paper presents a preliminary explainability-based analysis of three representative AI-generated image detectors, covering frequency-domain, embedding-based, and foundation-model-based approaches. Experiments are conducted on images generated by Generative Adversarial Networks (GANs) and diffusion models across multiple semantic categories. Post-hoc explainability techniques are employed to investigate the spatial and spectral regions driving detector decisions. The results reveal distinct and complementary behaviours among the considered methods, highlighting differences in sensitivity to semantic content, frequency patterns, and generation artifacts. In addition, the impact of lightweight dissemination transformations, such as image compression, is analysed through their effects on detector explanations. Rather than introducing a new detection model, this work aims to improve the understanding of existing AI-generated image detectors by framing explainability as a tool for forensic analysis and robustness assessment.

Explainability-Driven Analysis of AI-Generated Image Detectors: Insights into Generalization and Failure Modes

Pero, Chiara;
2026-01-01

Abstract

The increasing realism of AI-generated images has raised critical concerns about media authenticity and the reliability of automated detection systems. Although numerous detectors have been proposed, limited effort has been devoted to analysing the cues they exploit and the factors influencing their generalization behaviour. This paper presents a preliminary explainability-based analysis of three representative AI-generated image detectors, covering frequency-domain, embedding-based, and foundation-model-based approaches. Experiments are conducted on images generated by Generative Adversarial Networks (GANs) and diffusion models across multiple semantic categories. Post-hoc explainability techniques are employed to investigate the spatial and spectral regions driving detector decisions. The results reveal distinct and complementary behaviours among the considered methods, highlighting differences in sensitivity to semantic content, frequency patterns, and generation artifacts. In addition, the impact of lightweight dissemination transformations, such as image compression, is analysed through their effects on detector explanations. Rather than introducing a new detection model, this work aims to improve the understanding of existing AI-generated image detectors by framing explainability as a tool for forensic analysis and robustness assessment.
2026
9783032233462
9783032233479
AI-Generated Image Detection
Explainable AI
Media Forensics
Robustness Analysis
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14085/68541
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact