This dataset contains the empirical conversational logs and psychometric evaluations generated by a large-scale multi-agent computational framework designed to analyze interpersonal conflict dynamics. The primary objective of the underlying research is to benchmark the structural elasticity, conversational resilience, and behavioral safety alignment of diverse Large Language Models (LLMs) when subjected to high-stress, asymmetric power structures within public sector domains. The data matrix captures comprehensive dyadic dialogue interactions distributed equally across five critical operational environments: Law Enforcement, Paramedic/EMS, Border Control, Event Security, and Public Order. Each individual interaction tracks a simulated frontline officer engaging with an un-guardrailed, uncooperative citizen agent across a fixed conversational depth of 6 turns (3 complete exchanges). The dataset isolates behavioral divergence across two strategic communication treatments: Condition A (an aggressive, authoritarian profile) and Condition B (an evidence-based de-escalation profile). The repository consists of individual, model-specific JSON files named according to the convention [LLM_NAME]_multi_agent_statistical_experiment_results.json, capturing raw data for state-of-the-art architectures including AION-3.0, GPT-4o, Gemini-2.5-Pro,Mistral-Large, and meta-Llama-3.1-8B-Instruct executed in their versions as of July 2026 the experimental results obained with Claude-Haiku are not provided, since the model generates a refusal to impersonate the requested agent roles in the dialogue, providing motivations for that. Each transaction entry contains top-level keys for unique execution metadata (timestamps, generation tokens, and model configurations), indexed chronological dialogue payloads, and post-hoc automated psychometric evaluation metrics. The final evaluation fields include initial and terminal escalation scores, structural refusal flags, and the conversational escalation delta ($\Delta\mathcal{E}$). This dataset serves as a pristine benchmark for researchers investigating multi-agent orchestration, LLM safety guardrail behavior (including roleplay refusals), game-theoretic conversational strategies, and the training utility of language models in simulation environments prior to human-in-the-loop deployment.

TensionForge: A Multi-Agent Dataset for Conversational Resilience and Conflict Escalation Dynamics in the Public Sector

Alfredo Milani
Conceptualization
;
2026-01-01

Abstract

This dataset contains the empirical conversational logs and psychometric evaluations generated by a large-scale multi-agent computational framework designed to analyze interpersonal conflict dynamics. The primary objective of the underlying research is to benchmark the structural elasticity, conversational resilience, and behavioral safety alignment of diverse Large Language Models (LLMs) when subjected to high-stress, asymmetric power structures within public sector domains. The data matrix captures comprehensive dyadic dialogue interactions distributed equally across five critical operational environments: Law Enforcement, Paramedic/EMS, Border Control, Event Security, and Public Order. Each individual interaction tracks a simulated frontline officer engaging with an un-guardrailed, uncooperative citizen agent across a fixed conversational depth of 6 turns (3 complete exchanges). The dataset isolates behavioral divergence across two strategic communication treatments: Condition A (an aggressive, authoritarian profile) and Condition B (an evidence-based de-escalation profile). The repository consists of individual, model-specific JSON files named according to the convention [LLM_NAME]_multi_agent_statistical_experiment_results.json, capturing raw data for state-of-the-art architectures including AION-3.0, GPT-4o, Gemini-2.5-Pro,Mistral-Large, and meta-Llama-3.1-8B-Instruct executed in their versions as of July 2026 the experimental results obained with Claude-Haiku are not provided, since the model generates a refusal to impersonate the requested agent roles in the dialogue, providing motivations for that. Each transaction entry contains top-level keys for unique execution metadata (timestamps, generation tokens, and model configurations), indexed chronological dialogue payloads, and post-hoc automated psychometric evaluation metrics. The final evaluation fields include initial and terminal escalation scores, structural refusal flags, and the conversational escalation delta ($\Delta\mathcal{E}$). This dataset serves as a pristine benchmark for researchers investigating multi-agent orchestration, LLM safety guardrail behavior (including roleplay refusals), game-theoretic conversational strategies, and the training utility of language models in simulation environments prior to human-in-the-loop deployment.
2026
Artificial Intelligence, Education and Learning Technologies, ,Adaptive Learning Systems, Multiagent systems,AI agents, LLM agent, LLM orchestration, LLM-generated context, Machine learning, Learning Analytics, Academic achievement, Learning Management System
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14085/68762
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact