TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Jiaxuan, Ouyang, Xuan, Chen, Zhiyu, Hu, Yulan, Pan, Zheng, Li, Xin, Guo, Lan-Zhe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
por: Wang, Jiaxuan, et al.
Publicado: (2026)
por: Wang, Jiaxuan, et al.
Publicado: (2026)
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
por: Chen, Ge, et al.
Publicado: (2024)
por: Chen, Ge, et al.
Publicado: (2024)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
por: Li, Gengsheng, et al.
Publicado: (2026)
por: Li, Gengsheng, et al.
Publicado: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
por: Zhang, Songming, et al.
Publicado: (2025)
por: Zhang, Songming, et al.
Publicado: (2025)
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
por: Zheng, Lulu, et al.
Publicado: (2026)
por: Zheng, Lulu, et al.
Publicado: (2026)
TRACE: Temporal Routing with Autoregressive Cross-channel Experts for EEG Representation Learning
por: Ma, Fan, et al.
Publicado: (2026)
por: Ma, Fan, et al.
Publicado: (2026)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
por: Hu, Yulan, et al.
Publicado: (2023)
por: Hu, Yulan, et al.
Publicado: (2023)
TRACE-Bot: Detecting Emerging LLM-Driven Social Bots via Implicit Semantic Representations and AIGC-Enhanced Behavioral Patterns
por: Wang, Zhongbo, et al.
Publicado: (2026)
por: Wang, Zhongbo, et al.
Publicado: (2026)
TIP: Token Importance in On-Policy Distillation
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
TRACE: Learning to Compute on Circuit Graphs
por: Zheng, Ziyang, et al.
Publicado: (2025)
por: Zheng, Ziyang, et al.
Publicado: (2025)
Reasoning Fails Where Step Flow Breaks
por: Xu, Xiaoyu, et al.
Publicado: (2026)
por: Xu, Xiaoyu, et al.
Publicado: (2026)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
por: Liu, Xiaogeng, et al.
Publicado: (2026)
por: Liu, Xiaogeng, et al.
Publicado: (2026)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
por: Hu, Yulan, et al.
Publicado: (2025)
por: Hu, Yulan, et al.
Publicado: (2025)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
por: Ding, Ken
Publicado: (2026)
por: Ding, Ken
Publicado: (2026)
Multilingual Safety Alignment via Self-Distillation
por: Qin, Ruiyang, et al.
Publicado: (2026)
por: Qin, Ruiyang, et al.
Publicado: (2026)
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
por: Li, Shaolong, et al.
Publicado: (2026)
por: Li, Shaolong, et al.
Publicado: (2026)
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
por: Yi, Hao, et al.
Publicado: (2026)
por: Yi, Hao, et al.
Publicado: (2026)
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
por: Yang, Ming, et al.
Publicado: (2026)
por: Yang, Ming, et al.
Publicado: (2026)
TRACE: Traceroute-based Internet Route change Analysis with Ensemble Learning
por: Suzuki, Raul, et al.
Publicado: (2026)
por: Suzuki, Raul, et al.
Publicado: (2026)
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
por: Wei, Zhenlin, et al.
Publicado: (2026)
por: Wei, Zhenlin, et al.
Publicado: (2026)
GUNDAM: Aligning Large Language Models with Graph Understanding
por: Ouyang, Sheng, et al.
Publicado: (2024)
por: Ouyang, Sheng, et al.
Publicado: (2024)
Self-Distillation for Multi-Token Prediction
por: Zhao, Guoliang, et al.
Publicado: (2026)
por: Zhao, Guoliang, et al.
Publicado: (2026)
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
por: Yin, Bo, et al.
Publicado: (2026)
por: Yin, Bo, et al.
Publicado: (2026)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
por: Zhang, Xinsen, et al.
Publicado: (2026)
por: Zhang, Xinsen, et al.
Publicado: (2026)
Mitigating Deceptive Alignment via Self-Monitoring
por: Ji, Jiaming, et al.
Publicado: (2025)
por: Ji, Jiaming, et al.
Publicado: (2025)
Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
por: Zhang, Xiaotian, et al.
Publicado: (2025)
por: Zhang, Xiaotian, et al.
Publicado: (2025)
TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology
por: Liu, Feng, et al.
Publicado: (2026)
por: Liu, Feng, et al.
Publicado: (2026)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
por: Hu, Shijing, et al.
Publicado: (2025)
por: Hu, Shijing, et al.
Publicado: (2025)
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
por: Hu, Yulan, et al.
Publicado: (2023)
por: Hu, Yulan, et al.
Publicado: (2023)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
por: Jia, Nan, et al.
Publicado: (2026)
por: Jia, Nan, et al.
Publicado: (2026)
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
por: Jiang, Yuxuan, et al.
Publicado: (2026)
por: Jiang, Yuxuan, et al.
Publicado: (2026)
Route Experts by Sequence, not by Token
por: Wen, Tiansheng, et al.
Publicado: (2025)
por: Wen, Tiansheng, et al.
Publicado: (2025)
TRACE: Capability-Targeted Agentic Training
por: Kang, Hangoo, et al.
Publicado: (2026)
por: Kang, Hangoo, et al.
Publicado: (2026)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
por: Ke, Junlong, et al.
Publicado: (2026)
por: Ke, Junlong, et al.
Publicado: (2026)
From Fact Overwriting to Knowledge Evolution: Causal Editing via On-Policy Self-Distillation
por: Li, Shuaike, et al.
Publicado: (2026)
por: Li, Shuaike, et al.
Publicado: (2026)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
por: Rai, Arushi, et al.
Publicado: (2026)
por: Rai, Arushi, et al.
Publicado: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
por: Krishnakumar, Arjun, et al.
Publicado: (2025)
por: Krishnakumar, Arjun, et al.
Publicado: (2025)
Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search
por: Cao, Zhiyu, et al.
Publicado: (2026)
por: Cao, Zhiyu, et al.
Publicado: (2026)
Ejemplares similares
-
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
por: Wang, Jiaxuan, et al.
Publicado: (2026) -
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
por: Chen, Ge, et al.
Publicado: (2024) -
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
por: Li, Gengsheng, et al.
Publicado: (2026) -
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
por: Zhang, Songming, et al.
Publicado: (2025) -
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
por: Zheng, Lulu, et al.
Publicado: (2026)