Large Language Models (LLMs) are increasingly being used to support scientific evaluation, despite being often prohibited by the publishers' policies, with concerns being raised about their resilience to adversarial manipulation. This paper examines the threat posed by such attacks in the academic peer review process. Through controlled experiments, we demonstrate that maliciously crafted papers can introduce biases in LLM-generated reviews, giving higher scores and providing less critical feedback to forged submissions. Our findings show that the integrity of AI-assisted reviewing systems can be compromised, with direct implications for the fairness and trustworthiness of scientific publishing. This study marks an initial step towards understanding this emerging risk and highlights the urgency for further research into securing LLM-supported peer review to find appropriate countermeasures.
Can I Bypass Reviewer 2? A Preliminary Study on LLMs Vulnerabilities for Scientific Reviews
Milani, Alfredo
;
2025-01-01
Abstract
Large Language Models (LLMs) are increasingly being used to support scientific evaluation, despite being often prohibited by the publishers' policies, with concerns being raised about their resilience to adversarial manipulation. This paper examines the threat posed by such attacks in the academic peer review process. Through controlled experiments, we demonstrate that maliciously crafted papers can introduce biases in LLM-generated reviews, giving higher scores and providing less critical feedback to forged submissions. Our findings show that the integrity of AI-assisted reviewing systems can be compromised, with direct implications for the fairness and trustworthiness of scientific publishing. This study marks an initial step towards understanding this emerging risk and highlights the urgency for further research into securing LLM-supported peer review to find appropriate countermeasures.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


