Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Dekun, Shi, Haochen, Sun, Zhiyuan, Liu, Bang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
by: Deng, Minghang, et al.
Published: (2025)
by: Deng, Minghang, et al.
Published: (2025)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
by: Bayarri-Planas, Jordi, et al.
Published: (2024)
by: Bayarri-Planas, Jordi, et al.
Published: (2024)
LLMs and the Human Condition
by: Wallis, Peter
Published: (2024)
by: Wallis, Peter
Published: (2024)
Discovering Differences in Strategic Behavior Between Humans and LLMs
by: Wang, Caroline, et al.
Published: (2026)
by: Wang, Caroline, et al.
Published: (2026)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
Social Learning through Interactions with Other Agents: A Survey
by: Hillier, Dylan, et al.
Published: (2024)
by: Hillier, Dylan, et al.
Published: (2024)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2025)
by: Nfissi, Alaa, et al.
Published: (2025)
PestMA: LLM-based Multi-Agent System for Informed Pest Management
by: Shi, Hongrui, et al.
Published: (2025)
by: Shi, Hongrui, et al.
Published: (2025)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
by: Khanna, Danush, et al.
Published: (2025)
by: Khanna, Danush, et al.
Published: (2025)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
Gyan: An Explainable Neuro-Symbolic Language Model
by: Srinivasan, Venkat, et al.
Published: (2026)
by: Srinivasan, Venkat, et al.
Published: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
by: Zhang, Luyan, et al.
Published: (2025)
by: Zhang, Luyan, et al.
Published: (2025)
Large Language Models (LLMs) for Requirements Engineering (RE): A Systematic Literature Review
by: Zadenoori, Mohammad Amin, et al.
Published: (2025)
by: Zadenoori, Mohammad Amin, et al.
Published: (2025)
LLMs Aren't Human: A Critical Perspective on LLM Personality
by: Zierahn, Kim, et al.
Published: (2026)
by: Zierahn, Kim, et al.
Published: (2026)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
by: Taskooh, Hamidreza Kazemi, et al.
Published: (2026)
by: Taskooh, Hamidreza Kazemi, et al.
Published: (2026)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
by: Adelson, Trevor, et al.
Published: (2026)
by: Adelson, Trevor, et al.
Published: (2026)
Knowledge-Aware Iterative Retrieval for Multi-Agent Systems
by: Song, Seyoung
Published: (2025)
by: Song, Seyoung
Published: (2025)
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
by: Venkatasubramanian, Venkat, et al.
Published: (2024)
by: Venkatasubramanian, Venkat, et al.
Published: (2024)
The Curious Case of In-Training Compression of State Space Models
by: Chahine, Makram, et al.
Published: (2025)
by: Chahine, Makram, et al.
Published: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
by: Teixeira, Tiago, et al.
Published: (2026)
by: Teixeira, Tiago, et al.
Published: (2026)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)
by: Hashemi, Helia, et al.
Published: (2024)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
by: Zhu, Qian, et al.
Published: (2026)
by: Zhu, Qian, et al.
Published: (2026)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates
by: Kaplanski, Pawel
Published: (2026)
by: Kaplanski, Pawel
Published: (2026)
MAP: Multi-user Personalization with Collaborative LLM-powered Agents
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
by: Saito, Kei
Published: (2026)
by: Saito, Kei
Published: (2026)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
by: Hildebrand, Samuel, et al.
Published: (2025)
by: Hildebrand, Samuel, et al.
Published: (2025)
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
by: Gill, Gurbinder, et al.
Published: (2025)
by: Gill, Gurbinder, et al.
Published: (2025)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
by: Altamirano-Gomez, Gerardo, et al.
Published: (2023)
by: Altamirano-Gomez, Gerardo, et al.
Published: (2023)
A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
MODP: Multi Objective Directional Prompting
by: Nema, Aashutosh, et al.
Published: (2025)
by: Nema, Aashutosh, et al.
Published: (2025)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
Lightweight LLMs for Network Attack Detection in IoT Networks
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
Open-TI: Open Traffic Intelligence with Augmented Language Model
by: Da, Longchao, et al.
Published: (2023)
by: Da, Longchao, et al.
Published: (2023)
Similar Items
-
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025) -
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
by: Deng, Minghang, et al.
Published: (2025) -
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
by: Williamson, Dane, et al.
Published: (2025) -
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
by: Bayarri-Planas, Jordi, et al.
Published: (2024) -
LLMs and the Human Condition
by: Wallis, Peter
Published: (2024)