TED: Training-Free Experience Distillation for Multimodal Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Shuozhi, Wang, Jinqing, Liu, Zihao, Yuan, Miaomiao, Peng, Haoran, Zhao, Jin, Wang, Bingwen, Wang, Haoyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MCTS-SQL: Light-Weight LLMs can Master the Text-to-SQL through Monte Carlo Tree Search
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2025)
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2025)
Topology-Aware Revival for Efficient Sparse Training
von: Jin, Meiling, et al.
Veröffentlicht: (2026)
von: Jin, Meiling, et al.
Veröffentlicht: (2026)
Validity-Calibrated Reasoning Distillation
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026)
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
Distilled Protein Backbone Generation
von: Xie, Liyang, et al.
Veröffentlicht: (2025)
von: Xie, Liyang, et al.
Veröffentlicht: (2025)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
von: Zhang, Wenjing, et al.
Veröffentlicht: (2026)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2026)
Automated Fusion of Multimodal Electronic Health Records for Better Medical Predictions
von: Cui, Suhan, et al.
Veröffentlicht: (2024)
von: Cui, Suhan, et al.
Veröffentlicht: (2024)
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
von: Wu, Xiaojun, et al.
Veröffentlicht: (2025)
von: Wu, Xiaojun, et al.
Veröffentlicht: (2025)
Annealing Self-Distillation Rectification Improves Adversarial Training
von: Wu, Yu-Yu, et al.
Veröffentlicht: (2023)
von: Wu, Yu-Yu, et al.
Veröffentlicht: (2023)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
LLM and GNN are Complementary: Distilling LLM for Multimodal Graph Learning
von: Xu, Junjie, et al.
Veröffentlicht: (2024)
von: Xu, Junjie, et al.
Veröffentlicht: (2024)
Training Free Guided Flow Matching with Optimal Control
von: Wang, Luran, et al.
Veröffentlicht: (2024)
von: Wang, Luran, et al.
Veröffentlicht: (2024)
Logit Distillation on Manifolds: Mapping by Learning
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
Training Multimodal Large Reasoning Models Needs Better Thoughts: A Three-Stage Framework for Long Chain-of-Thought Synthesis and Selection
von: Wang, Yizhi, et al.
Veröffentlicht: (2025)
von: Wang, Yizhi, et al.
Veröffentlicht: (2025)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
von: Li, Mengqi, et al.
Veröffentlicht: (2025)
von: Li, Mengqi, et al.
Veröffentlicht: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning
von: Tao, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Tao, Xiaoyu, et al.
Veröffentlicht: (2026)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
von: Yang, Wang, et al.
Veröffentlicht: (2026)
von: Yang, Wang, et al.
Veröffentlicht: (2026)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
PSNE: Efficient Spectral Sparsification Algorithms for Scaling Network Embedding
von: Lin, Longlong, et al.
Veröffentlicht: (2024)
von: Lin, Longlong, et al.
Veröffentlicht: (2024)
When and Why Adversarial Training Improves PINNs: A Neural Tangent Kernel Perspective
von: Cao, Yuan-dong, et al.
Veröffentlicht: (2026)
von: Cao, Yuan-dong, et al.
Veröffentlicht: (2026)
Medical Multimodal Foundation Models in Clinical Diagnosis and Treatment: Applications, Challenges, and Future Directions
von: Sun, Kai, et al.
Veröffentlicht: (2024)
von: Sun, Kai, et al.
Veröffentlicht: (2024)
Can Past Experience Accelerate LLM Reasoning?
von: Pan, Bo, et al.
Veröffentlicht: (2025)
von: Pan, Bo, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
R$^2$PO: Decoupling Training Trajectories from Inference Responses for LLM Reasoning
von: Wang, Jingchu, et al.
Veröffentlicht: (2026)
von: Wang, Jingchu, et al.
Veröffentlicht: (2026)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
von: Wu, Yecheng, et al.
Veröffentlicht: (2026)
von: Wu, Yecheng, et al.
Veröffentlicht: (2026)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
von: Wang, Shengjie, et al.
Veröffentlicht: (2026)
von: Wang, Shengjie, et al.
Veröffentlicht: (2026)
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
von: Zhang, Yanzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhi, et al.
Veröffentlicht: (2025)
Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformer
von: Wang, Jingya, et al.
Veröffentlicht: (2025)
von: Wang, Jingya, et al.
Veröffentlicht: (2025)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
von: Yu, Bowen, et al.
Veröffentlicht: (2026)
von: Yu, Bowen, et al.
Veröffentlicht: (2026)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
von: Zhang, Ruichen, et al.
Veröffentlicht: (2025)
von: Zhang, Ruichen, et al.
Veröffentlicht: (2025)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
TED: Turn Emphasis with Dialogue Feature Attention for Emotion Recognition in Conversation
von: Ono, Junya, et al.
Veröffentlicht: (2025)
von: Ono, Junya, et al.
Veröffentlicht: (2025)
SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training
von: Wang, Prince Zizhuang, et al.
Veröffentlicht: (2026)
von: Wang, Prince Zizhuang, et al.
Veröffentlicht: (2026)
DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
von: Sun, Haoxiang, et al.
Veröffentlicht: (2026)
von: Sun, Haoxiang, et al.
Veröffentlicht: (2026)
Data-Free Continual Learning of Server Models in Model-Heterogeneous Cloud-Device Collaboration
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
von: Wu, Jiahao, et al.
Veröffentlicht: (2026)
von: Wu, Jiahao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MCTS-SQL: Light-Weight LLMs can Master the Text-to-SQL through Monte Carlo Tree Search
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2025) -
Topology-Aware Revival for Efficient Sparse Training
von: Jin, Meiling, et al.
Veröffentlicht: (2026) -
Validity-Calibrated Reasoning Distillation
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026) -
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
von: Yang, Shidong, et al.
Veröffentlicht: (2026) -
Distilled Protein Backbone Generation
von: Xie, Liyang, et al.
Veröffentlicht: (2025)