Liliz-lab/llm-advers-eval: LLM Adversarial Evaluation (Zhang et al., 2025)

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Liliz-lab, harmonia-ml
Format: Recurso digital
Published: Zenodo 2025
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866902337372553216
author Liliz-lab
harmonia-ml
author_facet Liliz-lab
harmonia-ml
contents <p>This dataset accompanies the paper "Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities" (Zhang et al., 2025). It contains behavioral and simulated data from experiments evaluating the decision-making robustness of large language models (LLMs) under adversarial and dynamic conditions.</p> <p>The dataset includes results from two canonical paradigms:</p> <p>Two-Armed Bandit Task — tests exploration–exploitation balance across different models and decoding settings (e.g., temperature, top-p).</p> <p>Multi-Round Trust Task (MRTT) — examines cooperative and adaptive decision-making in social exchange between LLMs and adversarial agents.</p> <p>Each file records trial-level choices, rewards, and model parameters (e.g., temperature, top-p, model name), with aggregated performance metrics for human and model comparisons. The dataset supports reproducible behavioral analyses and provides a foundation for studying model-specific susceptibilities to manipulation, rigidity, and fairness recognition.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17398584
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Liliz-lab/llm-advers-eval: LLM Adversarial Evaluation (Zhang et al., 2025)
Liliz-lab
harmonia-ml
<p>This dataset accompanies the paper "Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities" (Zhang et al., 2025). It contains behavioral and simulated data from experiments evaluating the decision-making robustness of large language models (LLMs) under adversarial and dynamic conditions.</p> <p>The dataset includes results from two canonical paradigms:</p> <p>Two-Armed Bandit Task — tests exploration–exploitation balance across different models and decoding settings (e.g., temperature, top-p).</p> <p>Multi-Round Trust Task (MRTT) — examines cooperative and adaptive decision-making in social exchange between LLMs and adversarial agents.</p> <p>Each file records trial-level choices, rewards, and model parameters (e.g., temperature, top-p, model name), with aggregated performance metrics for human and model comparisons. The dataset supports reproducible behavioral analyses and provides a foundation for studying model-specific susceptibilities to manipulation, rigidity, and fairness recognition.</p>
title Liliz-lab/llm-advers-eval: LLM Adversarial Evaluation (Zhang et al., 2025)
url https://doi.org/10.5281/zenodo.17398584