CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Siyi, Shi, Jiajun, Ni, Shiwen, Zhang, Ge, Li, Shuaimin, Wang, Shijian, Wen, Zhoufutu, Li, Yizhi, Alinejad-Rokny, Hamid, Liu, Jiaheng, Yang, Min, Huang, Wenhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
von: Li, Shuaimin, et al.
Veröffentlicht: (2025)
von: Li, Shuaimin, et al.
Veröffentlicht: (2025)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
Small Language Model as Data Prospector for Large Language Model
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
Quantification of Large Language Model Distillation
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
von: Wang, Qiyao, et al.
Veröffentlicht: (2024)
von: Wang, Qiyao, et al.
Veröffentlicht: (2024)
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
von: Chen, Dingwei, et al.
Veröffentlicht: (2024)
von: Chen, Dingwei, et al.
Veröffentlicht: (2024)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
von: Chen, Dingwei, et al.
Veröffentlicht: (2025)
von: Chen, Dingwei, et al.
Veröffentlicht: (2025)
How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions
von: Liang, Yuheng, et al.
Veröffentlicht: (2024)
von: Liang, Yuheng, et al.
Veröffentlicht: (2024)
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
von: Chen, Guhong, et al.
Veröffentlicht: (2024)
von: Chen, Guhong, et al.
Veröffentlicht: (2024)
R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models
von: Zhang, Ziming, et al.
Veröffentlicht: (2026)
von: Zhang, Ziming, et al.
Veröffentlicht: (2026)
PersonaMath: Boosting Mathematical Reasoning via Persona-Driven Data Augmentation
von: Luo, Jing, et al.
Veröffentlicht: (2024)
von: Luo, Jing, et al.
Veröffentlicht: (2024)
Enhancing Monte Carlo Dropout Performance for Uncertainty Quantification
von: Asgharnezhad, Hamzeh, et al.
Veröffentlicht: (2025)
von: Asgharnezhad, Hamzeh, et al.
Veröffentlicht: (2025)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
von: Puerto, Haritz, et al.
Veröffentlicht: (2026)
von: Puerto, Haritz, et al.
Veröffentlicht: (2026)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
CRISP: Compressing Redundancy in Chain-of-Thought via Intrinsic Saliency Pruning
von: Lan, Yangsong, et al.
Veröffentlicht: (2026)
von: Lan, Yangsong, et al.
Veröffentlicht: (2026)
CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
von: Li, Haoxi, et al.
Veröffentlicht: (2025)
von: Li, Haoxi, et al.
Veröffentlicht: (2025)
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
ETAGE: Enhanced Test Time Adaptation with Integrated Entropy and Gradient Norms for Robust Model Performance
von: Shamsi, Afshar, et al.
Veröffentlicht: (2024)
von: Shamsi, Afshar, et al.
Veröffentlicht: (2024)
Advancing Medical Image Segmentation with Mini-Net: A Lightweight Solution Tailored for Efficient Segmentation of Medical Images
von: Javed, Syed, et al.
Veröffentlicht: (2024)
von: Javed, Syed, et al.
Veröffentlicht: (2024)
Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking
von: Sadeghi, Alireza, et al.
Veröffentlicht: (2025)
von: Sadeghi, Alireza, et al.
Veröffentlicht: (2025)
Training Superior Sparse Autoencoders for Instruct Models
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
von: Chai, Linzheng, et al.
Veröffentlicht: (2024)
von: Chai, Linzheng, et al.
Veröffentlicht: (2024)
Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
von: Li, Zirui, et al.
Veröffentlicht: (2026)
von: Li, Zirui, et al.
Veröffentlicht: (2026)
Graph-Based Chain-of-Thought Pruning for Reducing Redundant Reflections in Reasoning LLMs
von: Yuan, Hongyuan, et al.
Veröffentlicht: (2026)
von: Yuan, Hongyuan, et al.
Veröffentlicht: (2026)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
CLinNET: An Interpretable and Uncertainty‐Aware Deep Learning Framework for Multi‐Modal Clinical Genomics
von: Ivan Bakhshayeshi, et al.
Veröffentlicht: (2026)
von: Ivan Bakhshayeshi, et al.
Veröffentlicht: (2026)
ExtendAttack: Attacking Servers of LRMs via Extending Reasoning
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model Benchmarks
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
HiST: Histological Images Reconstruct Tumor Spatial Transcriptomics via MultiScale Fusion Deep Learning
von: Wei Li, et al.
Veröffentlicht: (2026)
von: Wei Li, et al.
Veröffentlicht: (2026)
AutoMV: An Automatic Multi-Agent System for Music Video Generation
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
A Diagnostic Model for Acute Lymphoblastic Leukemia Using Metaheuristics and Deep Learning Methods
von: Rahmani, Amir Masoud, et al.
Veröffentlicht: (2024)
von: Rahmani, Amir Masoud, et al.
Veröffentlicht: (2024)
SemanticST: Spatially Informed Semantic Graph Learning for Clustering, Integration, and Scalable Analysis of Spatial Transcriptomics
von: Zahedi, Roxana, et al.
Veröffentlicht: (2025)
von: Zahedi, Roxana, et al.
Veröffentlicht: (2025)
Focal Modulation and Bidirectional Feature Fusion Network for Medical Image Segmentation
von: Safdar, Moin, et al.
Veröffentlicht: (2025)
von: Safdar, Moin, et al.
Veröffentlicht: (2025)
Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals
von: Li, Hanze, et al.
Veröffentlicht: (2025)
von: Li, Hanze, et al.
Veröffentlicht: (2025)
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
von: Li, Shuaimin, et al.
Veröffentlicht: (2025) -
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025) -
Small Language Model as Data Prospector for Large Language Model
von: Ni, Shiwen, et al.
Veröffentlicht: (2024) -
Quantification of Large Language Model Distillation
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025) -
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
von: Wang, Qiyao, et al.
Veröffentlicht: (2024)