Video anomaly detection in multi-scenario surveillance is challenging because normal activity patterns vary dramatically across environments, making it harder to distinguish anomalies from legitimate activity. Existing weakly-supervised methods treat all scenarios with a single global model, ignoring this contextual prior. We propose SF-VAD, a scene-aware framework that enriches motion-centric I3D features with semantic scene context derived from a learned I3D-to-CLIP projection. Our approach (i) trains a projection network on a small set of paired I3D-CLIP features to produce pseudo-CLIP embeddings for every video, and (ii) injects the projected scene context additively into a temporal transformer encoder. Under the standard weakly-supervised protocol, SF-VAD achieves 88.27% frame-level AUC on the MSAD benchmark, closely approaching the multimodal pi-VAD (88.47%) while surpassing all existing I3D-based methods and requiring only two feature types compared to six modalities.

SF-VAD: Scene-Aware Video Anomaly Detection via Projected Semantic Context

Pero, Chiara
2026-01-01

Abstract

Video anomaly detection in multi-scenario surveillance is challenging because normal activity patterns vary dramatically across environments, making it harder to distinguish anomalies from legitimate activity. Existing weakly-supervised methods treat all scenarios with a single global model, ignoring this contextual prior. We propose SF-VAD, a scene-aware framework that enriches motion-centric I3D features with semantic scene context derived from a learned I3D-to-CLIP projection. Our approach (i) trains a projection network on a small set of paired I3D-CLIP features to produce pseudo-CLIP embeddings for every video, and (ii) injects the projected scene context additively into a temporal transformer encoder. Under the standard weakly-supervised protocol, SF-VAD achieves 88.27% frame-level AUC on the MSAD benchmark, closely approaching the multimodal pi-VAD (88.47%) while surpassing all existing I3D-based methods and requiring only two feature types compared to six modalities.
2026
feature projection
multi-scenario surveillance
scene awareness
video anomaly detection
weakly-supervised learning
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14085/68881
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact