Data Analysis and Performance Evaluation of Simulation Deduction Based on LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Shansi, Li, Min |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
por: Kim, Byungjun, et al.
Publicado: (2024)
por: Kim, Byungjun, et al.
Publicado: (2024)
The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
por: Lalai, Harsh Nishant, et al.
Publicado: (2025)
por: Lalai, Harsh Nishant, et al.
Publicado: (2025)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
por: Li, Hanyu, et al.
Publicado: (2025)
por: Li, Hanyu, et al.
Publicado: (2025)
Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
por: Bailis, Suma, et al.
Publicado: (2024)
por: Bailis, Suma, et al.
Publicado: (2024)
Evaluation Ethics of LLMs in Legal Domain
por: Zhang, Ruizhe, et al.
Publicado: (2024)
por: Zhang, Ruizhe, et al.
Publicado: (2024)
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
por: Ou, Jiao, et al.
Publicado: (2024)
por: Ou, Jiao, et al.
Publicado: (2024)
Revealing Algorithmic Deductive Circuits for Logical Reasoning
por: Nguyen, Phuong Minh, et al.
Publicado: (2026)
por: Nguyen, Phuong Minh, et al.
Publicado: (2026)
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
por: Mishra, Shubham, et al.
Publicado: (2025)
por: Mishra, Shubham, et al.
Publicado: (2025)
Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models
por: Li, Yitian, et al.
Publicado: (2024)
por: Li, Yitian, et al.
Publicado: (2024)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
por: Maltoni, Davide, et al.
Publicado: (2025)
por: Maltoni, Davide, et al.
Publicado: (2025)
Investigating the Robustness of Deductive Reasoning with Large Language Models
por: Hoppe, Fabian, et al.
Publicado: (2025)
por: Hoppe, Fabian, et al.
Publicado: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
por: Jing, Huihao, et al.
Publicado: (2025)
por: Jing, Huihao, et al.
Publicado: (2025)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
por: Pandey, Atharva, et al.
Publicado: (2025)
por: Pandey, Atharva, et al.
Publicado: (2025)
The Role of Deductive and Inductive Reasoning in Large Language Models
por: Cai, Chengkun, et al.
Publicado: (2024)
por: Cai, Chengkun, et al.
Publicado: (2024)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
por: Ansell, Rebecca, et al.
Publicado: (2026)
por: Ansell, Rebecca, et al.
Publicado: (2026)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
por: Zhan, Runzhe, et al.
Publicado: (2025)
por: Zhan, Runzhe, et al.
Publicado: (2025)
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
por: Li, Wu, et al.
Publicado: (2026)
por: Li, Wu, et al.
Publicado: (2026)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
por: Tang, Shuo, et al.
Publicado: (2024)
por: Tang, Shuo, et al.
Publicado: (2024)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
por: Chen, Michael K., et al.
Publicado: (2025)
por: Chen, Michael K., et al.
Publicado: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
por: Lunardi, Riccardo, et al.
Publicado: (2025)
por: Lunardi, Riccardo, et al.
Publicado: (2025)
IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction
por: He, Kaiyu, et al.
Publicado: (2024)
por: He, Kaiyu, et al.
Publicado: (2024)
Data-Augmentation-Based Dialectal Adaptation for LLMs
por: Faisal, Fahim, et al.
Publicado: (2024)
por: Faisal, Fahim, et al.
Publicado: (2024)
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
por: Yang, Yongjin, et al.
Publicado: (2024)
por: Yang, Yongjin, et al.
Publicado: (2024)
ItD: Large Language Models Can Teach Themselves Induction through Deduction
por: Sun, Wangtao, et al.
Publicado: (2024)
por: Sun, Wangtao, et al.
Publicado: (2024)
How Far Are We from Intelligent Visual Deductive Reasoning?
por: Zhang, Yizhe, et al.
Publicado: (2024)
por: Zhang, Yizhe, et al.
Publicado: (2024)
Unlocking Recursive Thinking of LLMs: Alignment via Refinement
por: Zhang, Haoke, et al.
Publicado: (2025)
por: Zhang, Haoke, et al.
Publicado: (2025)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
por: Guo, Yanzhu, et al.
Publicado: (2024)
por: Guo, Yanzhu, et al.
Publicado: (2024)
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
por: Schwager, Nils, et al.
Publicado: (2026)
por: Schwager, Nils, et al.
Publicado: (2026)
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
por: Abdaljalil, Samir, et al.
Publicado: (2025)
por: Abdaljalil, Samir, et al.
Publicado: (2025)
Project SHADOW: Symbolic Higher-order Associative Deductive reasoning On Wikidata using LM probing
por: Akl, Hanna Abi
Publicado: (2024)
por: Akl, Hanna Abi
Publicado: (2024)
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
por: Shu, Fan, et al.
Publicado: (2026)
por: Shu, Fan, et al.
Publicado: (2026)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
por: Liang, Yiming, et al.
Publicado: (2025)
por: Liang, Yiming, et al.
Publicado: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
por: Jiang, Botian, et al.
Publicado: (2024)
por: Jiang, Botian, et al.
Publicado: (2024)
SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
por: Wu, Mengsong, et al.
Publicado: (2025)
por: Wu, Mengsong, et al.
Publicado: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
MoralBench: Moral Evaluation of LLMs
por: Ji, Jianchao, et al.
Publicado: (2024)
por: Ji, Jianchao, et al.
Publicado: (2024)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
por: Balloccu, Simone, et al.
Publicado: (2024)
por: Balloccu, Simone, et al.
Publicado: (2024)
From Guidelines to Guarantees: A Graph-Based Evaluation Harness for Domain-Specific Evaluation of LLMs
por: Lundin, Jessica M., et al.
Publicado: (2025)
por: Lundin, Jessica M., et al.
Publicado: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
por: Zhang, Ran, et al.
Publicado: (2024)
por: Zhang, Ran, et al.
Publicado: (2024)
Ejemplares similares
-
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
por: Kim, Byungjun, et al.
Publicado: (2024) -
The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
por: Lalai, Harsh Nishant, et al.
Publicado: (2025) -
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
por: Li, Hanyu, et al.
Publicado: (2025) -
Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
por: Bailis, Suma, et al.
Publicado: (2024) -
Evaluation Ethics of LLMs in Legal Domain
por: Zhang, Ruizhe, et al.
Publicado: (2024)