Large Language Models (LLMs) are increasingly being used to support scientific evaluation, despite being often prohibited by the publishers' policies, with concerns being raised about their resilience to adversarial manipulation. This paper examines the threat posed by such attacks in the academic peer review process. Through controlled experiments, we demonstrate that maliciously crafted papers can introduce biases in LLM-generated reviews, giving higher scores and providing less critical feedback to forged submissions. Our findings show that the integrity of AI-assisted reviewing systems can be compromised, with direct implications for the fairness and trustworthiness of scientific publishing. This study marks an initial step towards understanding this emerging risk and highlights the urgency for further research into securing LLM-supported peer review to find appropriate countermeasures.

Can I Bypass Reviewer 2? A Preliminary Study on LLMs Vulnerabilities for Scientific Reviews

Milani, Alfredo
;
2025-01-01

Abstract

Large Language Models (LLMs) are increasingly being used to support scientific evaluation, despite being often prohibited by the publishers' policies, with concerns being raised about their resilience to adversarial manipulation. This paper examines the threat posed by such attacks in the academic peer review process. Through controlled experiments, we demonstrate that maliciously crafted papers can introduce biases in LLM-generated reviews, giving higher scores and providing less critical feedback to forged submissions. Our findings show that the integrity of AI-assisted reviewing systems can be compromised, with direct implications for the fairness and trustworthiness of scientific publishing. This study marks an initial step towards understanding this emerging risk and highlights the urgency for further research into securing LLM-supported peer review to find appropriate countermeasures.
2025
artificial intelligence
cybersecurity
ethics
LLM
prompt injection
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14085/70041
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact