Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Zizhuo, Liu, Quanling, Quan, Jinsheng, Zhang, Chao, Zhu, Yifan, Shi, Xing, Xu, Jingtao, Li, Zhihui, Luo, Yawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
di: Quan, Jinsheng, et al.
Pubblicazione: (2025)
di: Quan, Jinsheng, et al.
Pubblicazione: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
di: Zhang, Xueqiao, et al.
Pubblicazione: (2025)
di: Zhang, Xueqiao, et al.
Pubblicazione: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
di: Li, Yubo, et al.
Pubblicazione: (2026)
di: Li, Yubo, et al.
Pubblicazione: (2026)
Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
di: van Sprang, Angela, et al.
Pubblicazione: (2025)
di: van Sprang, Angela, et al.
Pubblicazione: (2025)
MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures -- A Comprehensive Framework
di: Zhu, Yifan, et al.
Pubblicazione: (2025)
di: Zhu, Yifan, et al.
Pubblicazione: (2025)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
di: He, Jiahang, et al.
Pubblicazione: (2025)
di: He, Jiahang, et al.
Pubblicazione: (2025)
Turn, Turn, Turn...but Still Finding the Answers.
di: Boardman, Edna M.
Pubblicazione: (1994)
di: Boardman, Edna M.
Pubblicazione: (1994)
A Partially Observed Stochastic Linear Stackelberg Differential Game with Poisson Jumps under Mean-Variance Criteria
di: Lin, Jingtao, et al.
Pubblicazione: (2026)
di: Lin, Jingtao, et al.
Pubblicazione: (2026)
Global Maximum Principle for Partially Observed Risk-Sensitive Progressive Optimal Control of FBSDE with Poisson Jumps
di: Lin, Jingtao, et al.
Pubblicazione: (2025)
di: Lin, Jingtao, et al.
Pubblicazione: (2025)
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model
di: Wang, Chunshi, et al.
Pubblicazione: (2025)
di: Wang, Chunshi, et al.
Pubblicazione: (2025)
DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
di: Zhang, Chao, et al.
Pubblicazione: (2025)
di: Zhang, Chao, et al.
Pubblicazione: (2025)
On-Policy Context Distillation for Language Models
di: Ye, Tianzhu, et al.
Pubblicazione: (2026)
di: Ye, Tianzhu, et al.
Pubblicazione: (2026)
Exact Eigenvalues and Eigenvectors for Some n-Dimensional Matrices
di: Deng, Quanling
Pubblicazione: (2024)
di: Deng, Quanling
Pubblicazione: (2024)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
di: Zhao, Siyan, et al.
Pubblicazione: (2026)
di: Zhao, Siyan, et al.
Pubblicazione: (2026)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
di: Li, Songze, et al.
Pubblicazione: (2026)
di: Li, Songze, et al.
Pubblicazione: (2026)
Optimal Partition for Multi-Type Queueing System
di: Cao, Shengyu, et al.
Pubblicazione: (2021)
di: Cao, Shengyu, et al.
Pubblicazione: (2021)
Context-aware Difference Distilling for Multi-change Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Optimized View and Geometry Distillation from Multi-view Diffuser
di: Zhang, Youjia, et al.
Pubblicazione: (2023)
di: Zhang, Youjia, et al.
Pubblicazione: (2023)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
di: Lazaridis, Aristotelis, et al.
Pubblicazione: (2026)
di: Lazaridis, Aristotelis, et al.
Pubblicazione: (2026)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
di: Yan, Jianxin, et al.
Pubblicazione: (2025)
di: Yan, Jianxin, et al.
Pubblicazione: (2025)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
di: Zhang, Ruochen, et al.
Pubblicazione: (2024)
di: Zhang, Ruochen, et al.
Pubblicazione: (2024)
Same Same, But Different: An Examination of Different Student Groups' Information Behaviors
di: Rashika Bahl, et al.
Pubblicazione: (2025)
di: Rashika Bahl, et al.
Pubblicazione: (2025)
Same Image, Different Meanings: Toward Retrieval of Context-Dependent Meanings
di: Tsutsumi, Ayuto, et al.
Pubblicazione: (2026)
di: Tsutsumi, Ayuto, et al.
Pubblicazione: (2026)
ACR: Adaptive Context Refactoring via Context Refactoring Operators for Multi-Turn Dialogue
di: Shen, Jiawei, et al.
Pubblicazione: (2026)
di: Shen, Jiawei, et al.
Pubblicazione: (2026)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
di: Jiang, Yuhua, et al.
Pubblicazione: (2025)
di: Jiang, Yuhua, et al.
Pubblicazione: (2025)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
di: Zhang, Xinsen, et al.
Pubblicazione: (2026)
di: Zhang, Xinsen, et al.
Pubblicazione: (2026)
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
di: Wang, Zhebo, et al.
Pubblicazione: (2026)
di: Wang, Zhebo, et al.
Pubblicazione: (2026)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
di: Zou, Bo, et al.
Pubblicazione: (2024)
di: Zou, Bo, et al.
Pubblicazione: (2024)
Method for ZVS Implementation in High‐Voltage Arbitrary Waveform Generators Based on CHB Topology
di: Fuchao Lu, et al.
Pubblicazione: (2025)
di: Fuchao Lu, et al.
Pubblicazione: (2025)
Design of a 21‐Level Wide‐Frequency Range High‐Voltage Arbitrary Waveform Generator
di: Fuchao Lu, et al.
Pubblicazione: (2025)
di: Fuchao Lu, et al.
Pubblicazione: (2025)
An Eulerian Data Assimilation Method for Two-Layer Quasi-Geostrophic Model in Physical Domain
di: Yun, Hyeonggeun, et al.
Pubblicazione: (2025)
di: Yun, Hyeonggeun, et al.
Pubblicazione: (2025)
Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding
di: Lin, Zhiming
Pubblicazione: (2025)
di: Lin, Zhiming
Pubblicazione: (2025)
Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment
di: Cai, Yuang, et al.
Pubblicazione: (2024)
di: Cai, Yuang, et al.
Pubblicazione: (2024)
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
di: Lin, Xingwei, et al.
Pubblicazione: (2026)
di: Lin, Xingwei, et al.
Pubblicazione: (2026)
Three-Level Multi-Leader-Follower Incentive Stackelberg Differential Game with $H_\infty$ Constraint
di: Xiang, Na, et al.
Pubblicazione: (2024)
di: Xiang, Na, et al.
Pubblicazione: (2024)
Advances in 4D Generation: A Survey
di: Miao, Qiaowei, et al.
Pubblicazione: (2025)
di: Miao, Qiaowei, et al.
Pubblicazione: (2025)
Same Same But Different: Preventing Refactoring Attacks on Software Plagiarism Detection
di: Maisch, Robin, et al.
Pubblicazione: (2025)
di: Maisch, Robin, et al.
Pubblicazione: (2025)
Fertility fibres and coproduct coefficients in the LOT Hopf algebra
di: Zhu, Zhicheng, et al.
Pubblicazione: (2026)
di: Zhu, Zhicheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
di: Quan, Jinsheng, et al.
Pubblicazione: (2025) -
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
di: Zhang, Xueqiao, et al.
Pubblicazione: (2025) -
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026) -
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
di: Li, Yubo, et al.
Pubblicazione: (2026) -
Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
di: van Sprang, Angela, et al.
Pubblicazione: (2025)