Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Junyu, Fang, Tianqing, Zhang, Zhisong, Zhang, Hongming, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
von: Hu, Minda, et al.
Veröffentlicht: (2025)
von: Hu, Minda, et al.
Veröffentlicht: (2025)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
von: Deng, Chenlong, et al.
Veröffentlicht: (2025)
von: Deng, Chenlong, et al.
Veröffentlicht: (2025)
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
von: Li, Shuaiyi, et al.
Veröffentlicht: (2025)
von: Li, Shuaiyi, et al.
Veröffentlicht: (2025)
WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
von: Zhang, Zhisong, et al.
Veröffentlicht: (2024)
von: Zhang, Zhisong, et al.
Veröffentlicht: (2024)
Abstraction-of-Thought Makes Language Models Better Reasoners
von: Hong, Ruixin, et al.
Veröffentlicht: (2024)
von: Hong, Ruixin, et al.
Veröffentlicht: (2024)
Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
von: Chen, Xinghao, et al.
Veröffentlicht: (2025)
von: Chen, Xinghao, et al.
Veröffentlicht: (2025)
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
von: Das, Souvik, et al.
Veröffentlicht: (2024)
von: Das, Souvik, et al.
Veröffentlicht: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2026)
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2026)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
von: Li, Miao, et al.
Veröffentlicht: (2026)
von: Li, Miao, et al.
Veröffentlicht: (2026)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
von: Shangguan, Haonan, et al.
Veröffentlicht: (2025)
von: Shangguan, Haonan, et al.
Veröffentlicht: (2025)
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
von: He, Yancheng, et al.
Veröffentlicht: (2025)
von: He, Yancheng, et al.
Veröffentlicht: (2025)
Chain of Thought with Explicit Evidence Reasoning for Few-shot Relation Extraction
von: Ma, Xilai, et al.
Veröffentlicht: (2023)
von: Ma, Xilai, et al.
Veröffentlicht: (2023)
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
von: Liu, Zhenghao, et al.
Veröffentlicht: (2026)
von: Liu, Zhenghao, et al.
Veröffentlicht: (2026)
Demystifying Long Chain-of-Thought Reasoning in LLMs
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2025)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2025)
Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
von: Liang, Tian, et al.
Veröffentlicht: (2025)
von: Liang, Tian, et al.
Veröffentlicht: (2025)
Keypoint-based Progressive Chain-of-Thought Distillation for LLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2024)
von: Feng, Kaituo, et al.
Veröffentlicht: (2024)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2025)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories
von: Lu, Sidi, et al.
Veröffentlicht: (2026)
von: Lu, Sidi, et al.
Veröffentlicht: (2026)
Verified Critical Step Optimization for LLM Agents
von: Li, Mukai, et al.
Veröffentlicht: (2026)
von: Li, Mukai, et al.
Veröffentlicht: (2026)
Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
von: Zhu, Dong-Hai, et al.
Veröffentlicht: (2025)
von: Zhu, Dong-Hai, et al.
Veröffentlicht: (2025)
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
von: Cao, Jie, et al.
Veröffentlicht: (2026)
von: Cao, Jie, et al.
Veröffentlicht: (2026)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
von: Jia, Mengzhao, et al.
Veröffentlicht: (2024)
von: Jia, Mengzhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
von: Hu, Minda, et al.
Veröffentlicht: (2025) -
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
von: Fang, Tianqing, et al.
Veröffentlicht: (2025) -
UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
von: Deng, Chenlong, et al.
Veröffentlicht: (2025) -
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025) -
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)