Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ruiqi, Wang, Lingxiang, Zheng, Hainan Zhang Zhiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025)
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025)
Privacy-Preserving Reasoning with Knowledge-Distilled Parametric Retrieval Augmented Generation
von: Chen, Jinwen, et al.
Veröffentlicht: (2025)
von: Chen, Jinwen, et al.
Veröffentlicht: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
von: Yan, Shaotian, et al.
Veröffentlicht: (2026)
von: Yan, Shaotian, et al.
Veröffentlicht: (2026)
Parameter Importance-Driven Continual Learning for Foundation Models
von: Wang, Lingxiang, et al.
Veröffentlicht: (2025)
von: Wang, Lingxiang, et al.
Veröffentlicht: (2025)
Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation
von: Zhang, Qianchi, et al.
Veröffentlicht: (2026)
von: Zhang, Qianchi, et al.
Veröffentlicht: (2026)
FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation
von: Lu, Shule, et al.
Veröffentlicht: (2025)
von: Lu, Shule, et al.
Veröffentlicht: (2025)
Enhancing Chain of Thought Prompting in Large Language Models via Reasoning Patterns
von: Zhang, Yufeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yufeng, et al.
Veröffentlicht: (2024)
AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models
von: Zhang, Qianchi, et al.
Veröffentlicht: (2024)
von: Zhang, Qianchi, et al.
Veröffentlicht: (2024)
MaFeRw: Query Rewriting with Multi-Aspect Feedbacks for Retrieval-Augmented Large Language Models
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services
von: Chen, Jinwen, et al.
Veröffentlicht: (2026)
von: Chen, Jinwen, et al.
Veröffentlicht: (2026)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Learning to Erase Private Knowledge from Multi-Documents for Retrieval-Augmented Large Language Models
von: Wang, Yujing, et al.
Veröffentlicht: (2025)
von: Wang, Yujing, et al.
Veröffentlicht: (2025)
Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Models
von: Chen, Jinwen, et al.
Veröffentlicht: (2025)
von: Chen, Jinwen, et al.
Veröffentlicht: (2025)
Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation
von: Cui, Huizi, et al.
Veröffentlicht: (2026)
von: Cui, Huizi, et al.
Veröffentlicht: (2026)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation
von: Zhang, Lechen, et al.
Veröffentlicht: (2026)
von: Zhang, Lechen, et al.
Veröffentlicht: (2026)
TokAlign: Efficient Vocabulary Adaptation via Token Alignment
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
von: He, Yinhan, et al.
Veröffentlicht: (2025)
von: He, Yinhan, et al.
Veröffentlicht: (2025)
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
von: Kim, Minsang, et al.
Veröffentlicht: (2026)
von: Kim, Minsang, et al.
Veröffentlicht: (2026)
TwT: Thinking without Tokens by Habitual Reasoning Distillation with Multi-Teachers' Guidance
von: Xu, Jingxian, et al.
Veröffentlicht: (2025)
von: Xu, Jingxian, et al.
Veröffentlicht: (2025)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment
von: Li, Chong, et al.
Veröffentlicht: (2026)
von: Li, Chong, et al.
Veröffentlicht: (2026)
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
Continual Dialogue State Tracking via Reason-of-Select Distillation
von: Feng, Yujie, et al.
Veröffentlicht: (2024)
von: Feng, Yujie, et al.
Veröffentlicht: (2024)
GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics
von: Wang, Yujing, et al.
Veröffentlicht: (2026)
von: Wang, Yujing, et al.
Veröffentlicht: (2026)
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2026)
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2026)
CoT2Align: Cross-Chain of Thought Distillation via Optimal Transport Alignment for Language Models with Different Tokenizers
von: Le, Anh Duc, et al.
Veröffentlicht: (2025)
von: Le, Anh Duc, et al.
Veröffentlicht: (2025)
Self-Policy Distillation via Capability-Selective Subspace Projection
von: Hao, Guangya, et al.
Veröffentlicht: (2026)
von: Hao, Guangya, et al.
Veröffentlicht: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
von: Lin, Liang, et al.
Veröffentlicht: (2026)
von: Lin, Liang, et al.
Veröffentlicht: (2026)
Critique-Guided Distillation for Robust Reasoning via Refinement
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026) -
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025) -
Privacy-Preserving Reasoning with Knowledge-Distilled Parametric Retrieval Augmented Generation
von: Chen, Jinwen, et al.
Veröffentlicht: (2025) -
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
von: Yan, Shaotian, et al.
Veröffentlicht: (2026) -
Parameter Importance-Driven Continual Learning for Foundation Models
von: Wang, Lingxiang, et al.
Veröffentlicht: (2025)