Investigating CoT Monitorability in Large Reasoning Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Shu, Wu, Junchao, Gong, Xilin, Wu, Xuansheng, Wong, Derek, Liu, Ninghao, Wang, Di |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
di: Yang, Shu, et al.
Pubblicazione: (2025)
di: Yang, Shu, et al.
Pubblicazione: (2025)
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
di: Zhang, Jiayi, et al.
Pubblicazione: (2025)
di: Zhang, Jiayi, et al.
Pubblicazione: (2025)
MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
di: Wang, Yun, et al.
Pubblicazione: (2025)
di: Wang, Yun, et al.
Pubblicazione: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Rethinking Prompt-based Debiasing in Large Language Models
di: Yang, Xinyi, et al.
Pubblicazione: (2025)
di: Yang, Xinyi, et al.
Pubblicazione: (2025)
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning
di: Hu, Tianxiang, et al.
Pubblicazione: (2024)
di: Hu, Tianxiang, et al.
Pubblicazione: (2024)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Long or short CoT? Investigating Instance-level Switch of Large Reasoning Models
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
di: Li, Li, et al.
Pubblicazione: (2025)
di: Li, Li, et al.
Pubblicazione: (2025)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
di: Tong, Zhao, et al.
Pubblicazione: (2026)
di: Tong, Zhao, et al.
Pubblicazione: (2026)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
di: Gong, Xilin, et al.
Pubblicazione: (2026)
di: Gong, Xilin, et al.
Pubblicazione: (2026)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation
di: Li, Tong, et al.
Pubblicazione: (2025)
di: Li, Tong, et al.
Pubblicazione: (2025)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
di: Chen, Xin, et al.
Pubblicazione: (2026)
di: Chen, Xin, et al.
Pubblicazione: (2026)
Understanding Aha Moments: from External Observations to Internal Mechanisms
di: Yang, Shu, et al.
Pubblicazione: (2025)
di: Yang, Shu, et al.
Pubblicazione: (2025)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
di: Zhao, Shangziqi, et al.
Pubblicazione: (2025)
di: Zhao, Shangziqi, et al.
Pubblicazione: (2025)
SIM-CoT: Supervised Implicit Chain-of-Thought
di: Wei, Xilin, et al.
Pubblicazione: (2025)
di: Wei, Xilin, et al.
Pubblicazione: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
di: Chen, Jierun, et al.
Pubblicazione: (2025)
di: Chen, Jierun, et al.
Pubblicazione: (2025)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Efficient Long CoT Reasoning in Small Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
di: Yan, Shaotian, et al.
Pubblicazione: (2026)
di: Yan, Shaotian, et al.
Pubblicazione: (2026)
CoT-Driven Framework for Short Text Classification: Enhancing and Transferring Capabilities from Large to Smaller Model
di: Wu, Hui, et al.
Pubblicazione: (2024)
di: Wu, Hui, et al.
Pubblicazione: (2024)
Investigating Mysteries of CoT-Augmented Distillation
di: Wadhwa, Somin, et al.
Pubblicazione: (2024)
di: Wadhwa, Somin, et al.
Pubblicazione: (2024)
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
di: Jin, Senjie, et al.
Pubblicazione: (2025)
di: Jin, Senjie, et al.
Pubblicazione: (2025)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
di: Yang, Cehao, et al.
Pubblicazione: (2025)
di: Yang, Cehao, et al.
Pubblicazione: (2025)
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
di: Wu, Junchao, et al.
Pubblicazione: (2024)
di: Wu, Junchao, et al.
Pubblicazione: (2024)
Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements
di: Yang, Shu, et al.
Pubblicazione: (2025)
di: Yang, Shu, et al.
Pubblicazione: (2025)
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
di: Wu, Junchao, et al.
Pubblicazione: (2023)
di: Wu, Junchao, et al.
Pubblicazione: (2023)
Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
di: Wu, Mengsong, et al.
Pubblicazione: (2025)
di: Wu, Mengsong, et al.
Pubblicazione: (2025)
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
di: Li, Jiatong, et al.
Pubblicazione: (2025)
di: Li, Jiatong, et al.
Pubblicazione: (2025)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
di: Yang, Shu, et al.
Pubblicazione: (2025) -
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
di: Zhang, Jiayi, et al.
Pubblicazione: (2025) -
MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
di: Hu, Jingyu, et al.
Pubblicazione: (2025) -
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025) -
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)