Combining On-Policy Optimization and Distillation for Long-Context Reasoning in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramos, Miguel Moura, Alves, Duarte M., Martins, André F. T. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2025)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2025)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
von: Zhang, Xinsen, et al.
Veröffentlicht: (2026)
von: Zhang, Xinsen, et al.
Veröffentlicht: (2026)
On-Policy Context Distillation for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
Aligning Neural Machine Translation Models: Human Feedback in Training and Inference
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2023)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2023)
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
von: Palnitkar, Aadi, et al.
Veröffentlicht: (2026)
von: Palnitkar, Aadi, et al.
Veröffentlicht: (2026)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
von: Guan, Xin, et al.
Veröffentlicht: (2026)
von: Guan, Xin, et al.
Veröffentlicht: (2026)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
von: Alves, Duarte M., et al.
Veröffentlicht: (2024)
von: Alves, Duarte M., et al.
Veröffentlicht: (2024)
E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning
von: Liao, Zihan, et al.
Veröffentlicht: (2024)
von: Liao, Zihan, et al.
Veröffentlicht: (2024)
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2024)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2024)
Long-Context Generalization with Sparse Attention
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation
von: Dong, Zican, et al.
Veröffentlicht: (2025)
von: Dong, Zican, et al.
Veröffentlicht: (2025)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
MiniLLM: On-Policy Distillation of Large Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2023)
von: Gu, Yuxian, et al.
Veröffentlicht: (2023)
Black-Box On-Policy Distillation of Large Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
Thus Spake Long-Context Large Language Model
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2025)
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2025)
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
von: Yan, Yuchen, et al.
Veröffentlicht: (2025)
von: Yan, Yuchen, et al.
Veröffentlicht: (2025)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
Making Long-Context Language Models Better Multi-Hop Reasoners
von: Li, Yanyang, et al.
Veröffentlicht: (2024)
von: Li, Yanyang, et al.
Veröffentlicht: (2024)
Knowledge Distillation for Temporal Knowledge Graph Reasoning with Large Language Models
von: Xing, Wang, et al.
Veröffentlicht: (2026)
von: Xing, Wang, et al.
Veröffentlicht: (2026)
Distilling Reasoning Ability from Large Language Models with Adaptive Thinking
von: Chen, Xiaoshu, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoshu, et al.
Veröffentlicht: (2024)
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
von: He, Muyu, et al.
Veröffentlicht: (2025)
von: He, Muyu, et al.
Veröffentlicht: (2025)
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
von: Wan, Fanqi, et al.
Veröffentlicht: (2025)
von: Wan, Fanqi, et al.
Veröffentlicht: (2025)
Foresight Optimization for Strategic Reasoning in Large Language Models
von: Wang, Jiashuo, et al.
Veröffentlicht: (2026)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2026)
Improving In-Context Learning with Reasoning Distillation
von: Sadeq, Nafis, et al.
Veröffentlicht: (2025)
von: Sadeq, Nafis, et al.
Veröffentlicht: (2025)
Out-of-Context Reasoning in Large Language Models
von: Shaki, Jonathan, et al.
Veröffentlicht: (2025)
von: Shaki, Jonathan, et al.
Veröffentlicht: (2025)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
Training-Free Long-Context Scaling of Large Language Models
von: An, Chenxin, et al.
Veröffentlicht: (2024)
von: An, Chenxin, et al.
Veröffentlicht: (2024)
Large Language Models are Limited in Out-of-Context Knowledge Reasoning
von: Hu, Peng, et al.
Veröffentlicht: (2024)
von: Hu, Peng, et al.
Veröffentlicht: (2024)
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
von: Chen, Longze, et al.
Veröffentlicht: (2024)
von: Chen, Longze, et al.
Veröffentlicht: (2024)
Should We Still Pretrain Encoders with Masked Language Modeling?
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2025)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2025)
LongEmotion: Measuring Emotional Intelligence of Large Language Models in Long-Context Interaction
von: Liu, Weichu, et al.
Veröffentlicht: (2025)
von: Liu, Weichu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2025) -
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
von: Zhang, Xinsen, et al.
Veröffentlicht: (2026) -
On-Policy Context Distillation for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026) -
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2026) -
Aligning Neural Machine Translation Models: Human Feedback in Training and Inference
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2023)