Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Ziyang, Yuan, Qingyue, Zhang, Linhai, Zhou, Deyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation
von: Zhang, Linhai, et al.
Veröffentlicht: (2025)
von: Zhang, Linhai, et al.
Veröffentlicht: (2025)
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
von: Shen, Zhenyi, et al.
Veröffentlicht: (2025)
von: Shen, Zhenyi, et al.
Veröffentlicht: (2025)
MA$^{2}$P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion
von: Zhang, Dingyi, et al.
Veröffentlicht: (2026)
von: Zhang, Dingyi, et al.
Veröffentlicht: (2026)
Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)
STAR: Constraint LoRA with Dynamic Active Learning for Data-Efficient Fine-Tuning of Large Language Models
von: Zhang, Linhai, et al.
Veröffentlicht: (2024)
von: Zhang, Linhai, et al.
Veröffentlicht: (2024)
Keypoint-based Progressive Chain-of-Thought Distillation for LLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2024)
von: Feng, Kaituo, et al.
Veröffentlicht: (2024)
PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation
von: Zhang, Linhai, et al.
Veröffentlicht: (2025)
von: Zhang, Linhai, et al.
Veröffentlicht: (2025)
Fine-grainedly Synthesize Streaming Data Based On Large Language Models With Graph Structure Understanding For Data Sparsity
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
DINER: Debiasing Aspect-based Sentiment Analysis with Multi-variable Causal Inference
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
Persuasion Should be Double-Blind: A Multi-Domain Dialogue Dataset With Faithfulness Based on Causal Theory of Mind
von: Zhang, Dingyi, et al.
Veröffentlicht: (2025)
von: Zhang, Dingyi, et al.
Veröffentlicht: (2025)
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs
von: Cao, Shidong, et al.
Veröffentlicht: (2026)
von: Cao, Shidong, et al.
Veröffentlicht: (2026)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
SynGraph: A Dynamic Graph-LLM Synthesis Framework for Sparse Streaming User Sentiment Modeling
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
von: Chen, Xinghao, et al.
Veröffentlicht: (2025)
von: Chen, Xinghao, et al.
Veröffentlicht: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026)
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026)
Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
CoT-Valve: Length-Compressible Chain-of-Thought Tuning
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
RGAR: Recurrence Generation-augmented Retrieval for Factual-aware Medical Question Answering
von: Liang, Sichu, et al.
Veröffentlicht: (2025)
von: Liang, Sichu, et al.
Veröffentlicht: (2025)
FedCoT: Federated Chain-of-Thought Distillation for Large Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
GCoT: Chain-of-Thought Prompt Learning for Graphs
von: Yu, Xingtong, et al.
Veröffentlicht: (2025)
von: Yu, Xingtong, et al.
Veröffentlicht: (2025)
Learning to Maximize Mutual Information for Chain-of-Thought Distillation
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2025)
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2025)
Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning
von: Quamar, Mohammad Atif, et al.
Veröffentlicht: (2025)
von: Quamar, Mohammad Atif, et al.
Veröffentlicht: (2025)
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2026)
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2026)
Student-in-the-Loop Chain-of-Thought Distillation via Generation-Time Selection
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL
von: Thaker, Khushboo, et al.
Veröffentlicht: (2025)
von: Thaker, Khushboo, et al.
Veröffentlicht: (2025)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
von: Xiong, Xuan, et al.
Veröffentlicht: (2026)
von: Xiong, Xuan, et al.
Veröffentlicht: (2026)
Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2025)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning Without Prompting
von: Wang, Xuezhi, et al.
Veröffentlicht: (2024)
von: Wang, Xuezhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
von: Ma, Ziyang, et al.
Veröffentlicht: (2025) -
Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation
von: Zhang, Linhai, et al.
Veröffentlicht: (2025) -
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
von: Shen, Zhenyi, et al.
Veröffentlicht: (2025) -
MA$^{2}$P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion
von: Zhang, Dingyi, et al.
Veröffentlicht: (2026) -
Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)