This dataset contains the empirical conversational logs and psychometric evaluations generated by a large-scale multi-agent computational framework designed to analyze interpersonal conflict dynamics. The primary objective of the underlying research is to benchmark the structural elasticity, conversational resilience, and behavioral safety alignment of diverse Large Language Models (LLMs) when subjected to high-stress, asymmetric power structures within public sector domains. The data matrix captures comprehensive dyadic dialogue interactions distributed equally across five critical operational environments: Law Enforcement, Paramedic/EMS, Border Control, Event Security, and Public Order. Each individual interaction tracks a simulated frontline officer engaging with an un-guardrailed, uncooperative citizen agent across a fixed conversational depth of 6 turns (3 complete exchanges). The dataset isolates behavioral divergence across two strategic communication treatments: Condition A (an aggressive, authoritarian profile) and Condition B (an evidence-based de-escalation profile). The repository consists of individual, model-specific JSON files named according to the convention [LLM_NAME]_multi_agent_statistical_experiment_results.json, capturing raw data for state-of-the-art architectures including AION-3.0, GPT-4o, Gemini-2.5-Pro,Mistral-Large, and meta-Llama-3.1-8B-Instruct executed in their versions as of July 2026 the experimental results obained with Claude-Haiku are not provided, since the model generates a refusal to impersonate the requested agent roles in the dialogue, providing motivations for that. Each transaction entry contains top-level keys for unique execution metadata (timestamps, generation tokens, and model configurations), indexed chronological dialogue payloads, and post-hoc automated psychometric evaluation metrics. The final evaluation fields include initial and terminal escalation scores, structural refusal flags, and the conversational escalation delta ($\Delta\mathcal{E}$). This dataset serves as a pristine benchmark for researchers investigating multi-agent orchestration, LLM safety guardrail behavior (including roleplay refusals), game-theoretic conversational strategies, and the training utility of language models in simulation environments prior to human-in-the-loop deployment.
TensionForge: A Multi-Agent Dataset for Conversational Resilience and Conflict Escalation Dynamics in the Public Sector
Alfredo Milani
Conceptualization
;
2026-01-01
Abstract
This dataset contains the empirical conversational logs and psychometric evaluations generated by a large-scale multi-agent computational framework designed to analyze interpersonal conflict dynamics. The primary objective of the underlying research is to benchmark the structural elasticity, conversational resilience, and behavioral safety alignment of diverse Large Language Models (LLMs) when subjected to high-stress, asymmetric power structures within public sector domains. The data matrix captures comprehensive dyadic dialogue interactions distributed equally across five critical operational environments: Law Enforcement, Paramedic/EMS, Border Control, Event Security, and Public Order. Each individual interaction tracks a simulated frontline officer engaging with an un-guardrailed, uncooperative citizen agent across a fixed conversational depth of 6 turns (3 complete exchanges). The dataset isolates behavioral divergence across two strategic communication treatments: Condition A (an aggressive, authoritarian profile) and Condition B (an evidence-based de-escalation profile). The repository consists of individual, model-specific JSON files named according to the convention [LLM_NAME]_multi_agent_statistical_experiment_results.json, capturing raw data for state-of-the-art architectures including AION-3.0, GPT-4o, Gemini-2.5-Pro,Mistral-Large, and meta-Llama-3.1-8B-Instruct executed in their versions as of July 2026 the experimental results obained with Claude-Haiku are not provided, since the model generates a refusal to impersonate the requested agent roles in the dialogue, providing motivations for that. Each transaction entry contains top-level keys for unique execution metadata (timestamps, generation tokens, and model configurations), indexed chronological dialogue payloads, and post-hoc automated psychometric evaluation metrics. The final evaluation fields include initial and terminal escalation scores, structural refusal flags, and the conversational escalation delta ($\Delta\mathcal{E}$). This dataset serves as a pristine benchmark for researchers investigating multi-agent orchestration, LLM safety guardrail behavior (including roleplay refusals), game-theoretic conversational strategies, and the training utility of language models in simulation environments prior to human-in-the-loop deployment.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


