Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Qiqi, Wang, Xinpeng, Mondorf, Philipp, Hedderich, Michael A., Plank, Barbara |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
por: Cheng, Xinyuan, et al.
Publicado: (2026)
por: Cheng, Xinyuan, et al.
Publicado: (2026)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
por: Orth, Jasmin, et al.
Publicado: (2025)
por: Orth, Jasmin, et al.
Publicado: (2025)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
por: Rabern, Brian, et al.
Publicado: (2026)
por: Rabern, Brian, et al.
Publicado: (2026)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
por: Bertolazzi, Leonardo, et al.
Publicado: (2025)
por: Bertolazzi, Leonardo, et al.
Publicado: (2025)
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
por: Ma, Bolei, et al.
Publicado: (2024)
por: Ma, Bolei, et al.
Publicado: (2024)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
por: Zhao, Raoyuan, et al.
Publicado: (2025)
por: Zhao, Raoyuan, et al.
Publicado: (2025)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
por: Eichin, Florian, et al.
Publicado: (2025)
por: Eichin, Florian, et al.
Publicado: (2025)
Reason to Rote: Rethinking Memorization in Reasoning
por: Du, Yupei, et al.
Publicado: (2025)
por: Du, Yupei, et al.
Publicado: (2025)
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
por: Wang, Xinpeng, et al.
Publicado: (2024)
por: Wang, Xinpeng, et al.
Publicado: (2024)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
por: Hedderich, Michael A., et al.
Publicado: (2025)
por: Hedderich, Michael A., et al.
Publicado: (2025)
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
por: Körner, Felicia, et al.
Publicado: (2026)
por: Körner, Felicia, et al.
Publicado: (2026)
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
por: Eichin, Florian, et al.
Publicado: (2025)
por: Eichin, Florian, et al.
Publicado: (2025)
Tracing Uncertainty in Language Model "Reasoning"
por: Grünefeld, Nils, et al.
Publicado: (2026)
por: Grünefeld, Nils, et al.
Publicado: (2026)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
por: Zhao, Raoyuan, et al.
Publicado: (2025)
por: Zhao, Raoyuan, et al.
Publicado: (2025)
Refusal Direction is Universal Across Safety-Aligned Languages
por: Wang, Xinpeng, et al.
Publicado: (2025)
por: Wang, Xinpeng, et al.
Publicado: (2025)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
por: Wang, Xinpeng, et al.
Publicado: (2024)
por: Wang, Xinpeng, et al.
Publicado: (2024)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
por: Chen, Beiduo, et al.
Publicado: (2024)
por: Chen, Beiduo, et al.
Publicado: (2024)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
por: Chen, Beiduo, et al.
Publicado: (2025)
por: Chen, Beiduo, et al.
Publicado: (2025)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
por: Wang, Yongjie, et al.
Publicado: (2025)
por: Wang, Yongjie, et al.
Publicado: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
por: Si, Shengyun, et al.
Publicado: (2025)
por: Si, Shengyun, et al.
Publicado: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
por: Wang, Xinpeng, et al.
Publicado: (2025)
por: Wang, Xinpeng, et al.
Publicado: (2025)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
por: Mondorf, Philipp, et al.
Publicado: (2025)
por: Mondorf, Philipp, et al.
Publicado: (2025)
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study
por: Ma, Bolei, et al.
Publicado: (2024)
por: Ma, Bolei, et al.
Publicado: (2024)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
por: Chen, Beiduo, et al.
Publicado: (2026)
por: Chen, Beiduo, et al.
Publicado: (2026)
When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training
por: Körner, Felicia, et al.
Publicado: (2026)
por: Körner, Felicia, et al.
Publicado: (2026)
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
por: Bai, Haolei, et al.
Publicado: (2026)
por: Bai, Haolei, et al.
Publicado: (2026)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
por: Liu, Yihong, et al.
Publicado: (2026)
por: Liu, Yihong, et al.
Publicado: (2026)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
por: Puerto, Haritz, et al.
Publicado: (2024)
por: Puerto, Haritz, et al.
Publicado: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
por: Wang, Xinpeng, et al.
Publicado: (2024)
por: Wang, Xinpeng, et al.
Publicado: (2024)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
por: Zuo, Longfei, et al.
Publicado: (2025)
por: Zuo, Longfei, et al.
Publicado: (2025)
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
por: Ahnert, Georg, et al.
Publicado: (2025)
por: Ahnert, Georg, et al.
Publicado: (2025)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
por: Zhao, Raoyuan, et al.
Publicado: (2026)
por: Zhao, Raoyuan, et al.
Publicado: (2026)
Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
por: Tao, Chaofan, et al.
Publicado: (2024)
por: Tao, Chaofan, et al.
Publicado: (2024)
Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling
por: Eichin, Florian, et al.
Publicado: (2024)
por: Eichin, Florian, et al.
Publicado: (2024)
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
por: Mondorf, Philipp, et al.
Publicado: (2025)
por: Mondorf, Philipp, et al.
Publicado: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
por: Wu, Yuyang, et al.
Publicado: (2025)
por: Wu, Yuyang, et al.
Publicado: (2025)
Ejemplares similares
-
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024) -
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
por: Cheng, Xinyuan, et al.
Publicado: (2026) -
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
por: Orth, Jasmin, et al.
Publicado: (2025) -
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
por: Mondorf, Philipp, et al.
Publicado: (2024) -
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
por: Mondorf, Philipp, et al.
Publicado: (2024)