GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Jinhao, Zhang, Renming, Diffenderfer, James, Kailkhura, Bhavya, Sun, Lichao, Stengel-Eskin, Elias, Bansal, Mohit, Chen, Tianlong, Xu, Kaidi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
von: Pothiraj, Atin, et al.
Veröffentlicht: (2025)
von: Pothiraj, Atin, et al.
Veröffentlicht: (2025)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
Teaching Models to Balance Resisting and Accepting Persuasion
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
Language Models Identify Ambiguities and Exploit Loopholes
von: Choi, Jio, et al.
Veröffentlicht: (2025)
von: Choi, Jio, et al.
Veröffentlicht: (2025)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
End-to-End Mesh Optimization of a Hybrid Deep Learning Black-Box PDE Solver
von: Ma, Shaocong, et al.
Veröffentlicht: (2024)
von: Ma, Shaocong, et al.
Veröffentlicht: (2024)
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
von: Hase, Peter, et al.
Veröffentlicht: (2024)
von: Hase, Peter, et al.
Veröffentlicht: (2024)
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
von: Nguyen, Duy, et al.
Veröffentlicht: (2024)
von: Nguyen, Duy, et al.
Veröffentlicht: (2024)
Soft Self-Consistency Improves Language Model Agents
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Are language models rational? The case of coherence norms and belief revision
von: Hofweber, Thomas, et al.
Veröffentlicht: (2024)
von: Hofweber, Thomas, et al.
Veröffentlicht: (2024)
AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
von: Khan, Zaid, et al.
Veröffentlicht: (2024)
von: Khan, Zaid, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Generation with Conflicting Evidence
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
von: Wan, David, et al.
Veröffentlicht: (2024)
von: Wan, David, et al.
Veröffentlicht: (2024)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
von: Ananthram, Amith, et al.
Veröffentlicht: (2024)
von: Ananthram, Amith, et al.
Veröffentlicht: (2024)
Multimodal Fact-Level Attribution for Verifiable Reasoning
von: Wan, David, et al.
Veröffentlicht: (2026)
von: Wan, David, et al.
Veröffentlicht: (2026)
MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
GenerationPrograms: Fine-grained Attribution with Executable Programs
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
von: Eliav, Ron, et al.
Veröffentlicht: (2025)
von: Eliav, Ron, et al.
Veröffentlicht: (2025)
Forecasting Fails: Unveiling Evasion Attacks in Weather Prediction Models
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2026)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2026)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
von: Duan, Jinhao, et al.
Veröffentlicht: (2025) -
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
von: Duan, Jinhao, et al.
Veröffentlicht: (2025) -
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023) -
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
von: Pothiraj, Atin, et al.
Veröffentlicht: (2025) -
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)