ADDI - Adversarial Data-Harvesting Defensive Injections Dataset This dataset is the data reference for the paper "Adversarial Defensive Injection: Prevention of Unauthorized RAG Data Harvesting via Multi-Vector Obfuscation" Authors: Alfredo Milani, Link Campus University of Rome, Emanuele Florindi, University of Modena and Reggio Emilia,Mattia Polticchia, University of Modena and Reggio Emili Reference Author: Alfredo Milani, [email protected] The core idea is to use indirect prompt injection as a defensive technique in order to prevent data harvesting from web site un-authorized from the data owner, an increasing phenomena with the advent of LLM in general and RAG systems. The DATASET File: dataset_adversarial_defense_injection.json is based on 80 news extracted from four categories of AP News data repository and used to generate different type of HTML pages: plain HTML pages used as control reference, and HTML containing defensive prompt injections to prevent specific LLM processing tasks, i,e,Categorization and Summarization, The HMTL propmpt injection are obfuscated and trasparent to a human user using four different obfuscation techniques: zero-sized, color-match, metadata, omni-obfuscation. The RESULTS File: Experiments_results_ALL_articles_injections_model_task.json contains the results obtained from the experiments executed on the previous HTML pages by submitting them to four widespread LLM (GPT, LLAMA, DEEPSEEK, GEMINI) requesting to execute the two different tasks, Categorization and Summarization.
ADDI - Adversarial Data-Harvesting Defensive Injections Dataset
Alfredo Milani
Conceptualization
;
2026-01-01
Abstract
ADDI - Adversarial Data-Harvesting Defensive Injections Dataset This dataset is the data reference for the paper "Adversarial Defensive Injection: Prevention of Unauthorized RAG Data Harvesting via Multi-Vector Obfuscation" Authors: Alfredo Milani, Link Campus University of Rome, Emanuele Florindi, University of Modena and Reggio Emilia,Mattia Polticchia, University of Modena and Reggio Emili Reference Author: Alfredo Milani, [email protected] The core idea is to use indirect prompt injection as a defensive technique in order to prevent data harvesting from web site un-authorized from the data owner, an increasing phenomena with the advent of LLM in general and RAG systems. The DATASET File: dataset_adversarial_defense_injection.json is based on 80 news extracted from four categories of AP News data repository and used to generate different type of HTML pages: plain HTML pages used as control reference, and HTML containing defensive prompt injections to prevent specific LLM processing tasks, i,e,Categorization and Summarization, The HMTL propmpt injection are obfuscated and trasparent to a human user using four different obfuscation techniques: zero-sized, color-match, metadata, omni-obfuscation. The RESULTS File: Experiments_results_ALL_articles_injections_model_task.json contains the results obtained from the experiments executed on the previous HTML pages by submitting them to four widespread LLM (GPT, LLAMA, DEEPSEEK, GEMINI) requesting to execute the two different tasks, Categorization and Summarization.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


