Guardado en:
| Autores principales: | Zhang, Yongheng, Chen, Qiguang, Zhou, Jingxuan, Wang, Peng, Si, Jiasheng, Wang, Jin, Lu, Wenpeng, Qin, Libo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2410.04463 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language Understanding
por: Qin, Libo, et al.
Publicado: (2024)
por: Qin, Libo, et al.
Publicado: (2024)
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
por: Chen, Qiguang, et al.
Publicado: (2024)
por: Chen, Qiguang, et al.
Publicado: (2024)
AutoCAP: Towards Automatic Cross-lingual Alignment Planning for Zero-shot Chain-of-Thought
por: Zhang, Yongheng, et al.
Publicado: (2024)
por: Zhang, Yongheng, et al.
Publicado: (2024)
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
por: Zhang, Yongheng, et al.
Publicado: (2025)
por: Zhang, Yongheng, et al.
Publicado: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
por: Chen, Qiguang, et al.
Publicado: (2025)
por: Chen, Qiguang, et al.
Publicado: (2025)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
por: Liu, Xu, et al.
Publicado: (2026)
por: Liu, Xu, et al.
Publicado: (2026)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
por: Zhang, Yongheng, et al.
Publicado: (2025)
por: Zhang, Yongheng, et al.
Publicado: (2025)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
por: Chen, Qiguang, et al.
Publicado: (2024)
por: Chen, Qiguang, et al.
Publicado: (2024)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
por: Yao, Jihan, et al.
Publicado: (2024)
por: Yao, Jihan, et al.
Publicado: (2024)
CHECKWHY: Causal Fact Verification via Argument Structure
por: Si, Jiasheng, et al.
Publicado: (2024)
por: Si, Jiasheng, et al.
Publicado: (2024)
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
por: Chen, Qiguang, et al.
Publicado: (2025)
por: Chen, Qiguang, et al.
Publicado: (2025)
DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective
por: Peng, Dengyun, et al.
Publicado: (2025)
por: Peng, Dengyun, et al.
Publicado: (2025)
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
por: Zhang, Chenyuan, et al.
Publicado: (2026)
por: Zhang, Chenyuan, et al.
Publicado: (2026)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
por: Cheng, Zihui, et al.
Publicado: (2025)
por: Cheng, Zihui, et al.
Publicado: (2025)
Beyond Surface Reasoning: Unveiling the True Long Chain-of-Thought Capacity of Diffusion Large Language Models
por: Chen, Qiguang, et al.
Publicado: (2025)
por: Chen, Qiguang, et al.
Publicado: (2025)
Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
por: Guan, Jiannan, et al.
Publicado: (2025)
por: Guan, Jiannan, et al.
Publicado: (2025)
Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis
por: Ling, Zipeng, et al.
Publicado: (2026)
por: Ling, Zipeng, et al.
Publicado: (2026)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
por: Wang, Peng, et al.
Publicado: (2025)
por: Wang, Peng, et al.
Publicado: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
por: Chen, Qiguang, et al.
Publicado: (2026)
por: Chen, Qiguang, et al.
Publicado: (2026)
SRLCG: Self-Rectified Large-Scale Code Generation with Multidimensional Chain-of-Thought and Dynamic Backtracking
por: Ma, Hongru, et al.
Publicado: (2025)
por: Ma, Hongru, et al.
Publicado: (2025)
Wrong as Sequence Violation: The Structural Definition of Wrong as Misordering Across Reasoning, Physics, Computation, Cognition, and Ethics
por: Stewart, Arthur
Publicado: (2026)
por: Stewart, Arthur
Publicado: (2026)
What is Wrong with Perplexity for Long-context Language Modeling?
por: Fang, Lizhe, et al.
Publicado: (2024)
por: Fang, Lizhe, et al.
Publicado: (2024)
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
por: Wang, Weiming, et al.
Publicado: (2026)
por: Wang, Weiming, et al.
Publicado: (2026)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
por: Shrestha, Robik, et al.
Publicado: (2020)
por: Shrestha, Robik, et al.
Publicado: (2020)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
por: Sun, Linzhuang, et al.
Publicado: (2025)
por: Sun, Linzhuang, et al.
Publicado: (2025)
Large Language Models Meet NLP: A Survey
por: Qin, Libo, et al.
Publicado: (2024)
por: Qin, Libo, et al.
Publicado: (2024)
Humans Perceive Wrong Narratives from AI Reasoning Texts
por: Levy, Mosh, et al.
Publicado: (2025)
por: Levy, Mosh, et al.
Publicado: (2025)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
por: Zhou, Shijia, et al.
Publicado: (2024)
por: Zhou, Shijia, et al.
Publicado: (2024)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
por: Balepur, Nishant, et al.
Publicado: (2023)
por: Balepur, Nishant, et al.
Publicado: (2023)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
por: Shafiei, Mohammadamin, et al.
Publicado: (2025)
por: Shafiei, Mohammadamin, et al.
Publicado: (2025)
Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
por: Chen, Qiguang, et al.
Publicado: (2025)
por: Chen, Qiguang, et al.
Publicado: (2025)
Easy Problems That LLMs Get Wrong
por: Williams, Sean, et al.
Publicado: (2024)
por: Williams, Sean, et al.
Publicado: (2024)
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
por: Su, Ruiran, et al.
Publicado: (2025)
por: Su, Ruiran, et al.
Publicado: (2025)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
por: Qin, Libo, et al.
Publicado: (2024)
por: Qin, Libo, et al.
Publicado: (2024)
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
por: Wen, Xueru, et al.
Publicado: (2024)
por: Wen, Xueru, et al.
Publicado: (2024)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
por: Si, Chenglei, et al.
Publicado: (2023)
por: Si, Chenglei, et al.
Publicado: (2023)
The Realignment Problem: When Right becomes Wrong in LLMs
por: Sharma, Aakash Sen, et al.
Publicado: (2025)
por: Sharma, Aakash Sen, et al.
Publicado: (2025)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
por: He, Yanjie
Publicado: (2026)
por: He, Yanjie
Publicado: (2026)
Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
por: Hagar, Nick, et al.
Publicado: (2025)
por: Hagar, Nick, et al.
Publicado: (2025)
What's Wrong? Refining Meeting Summaries with LLM Feedback
por: Kirstein, Frederic, et al.
Publicado: (2024)
por: Kirstein, Frederic, et al.
Publicado: (2024)
Ejemplares similares
-
CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language Understanding
por: Qin, Libo, et al.
Publicado: (2024) -
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
por: Chen, Qiguang, et al.
Publicado: (2024) -
AutoCAP: Towards Automatic Cross-lingual Alignment Planning for Zero-shot Chain-of-Thought
por: Zhang, Yongheng, et al.
Publicado: (2024) -
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
por: Zhang, Yongheng, et al.
Publicado: (2025) -
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
por: Chen, Qiguang, et al.
Publicado: (2025)