Rubric-based On-policy Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Fang, Junfeng, Hong, Zhepei, Zheng, Mao, Song, Mingyang, Li, Gengsheng, Jiang, Houcheng, Zhang, Dan, Guo, Haiyun, Wang, Xiang, Chua, Tat-Seng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
di: Fang, Junfeng, et al.
Pubblicazione: (2026)
di: Fang, Junfeng, et al.
Pubblicazione: (2026)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
di: Zhong, Qiyong, et al.
Pubblicazione: (2026)
di: Zhong, Qiyong, et al.
Pubblicazione: (2026)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
di: Fang, Junfeng, et al.
Pubblicazione: (2025)
di: Fang, Junfeng, et al.
Pubblicazione: (2025)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
di: Fang, Junfeng, et al.
Pubblicazione: (2025)
di: Fang, Junfeng, et al.
Pubblicazione: (2025)
AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models
di: Fang, Junfeng, et al.
Pubblicazione: (2024)
di: Fang, Junfeng, et al.
Pubblicazione: (2024)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
di: Sheng, Leheng, et al.
Pubblicazione: (2025)
di: Sheng, Leheng, et al.
Pubblicazione: (2025)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
di: Ye, Hao, et al.
Pubblicazione: (2026)
di: Ye, Hao, et al.
Pubblicazione: (2026)
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
di: Guo, Hongyu, et al.
Pubblicazione: (2025)
di: Guo, Hongyu, et al.
Pubblicazione: (2025)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
di: Hong, Zhepei, et al.
Pubblicazione: (2026)
di: Hong, Zhepei, et al.
Pubblicazione: (2026)
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
di: Jin, Zhe, et al.
Pubblicazione: (2025)
di: Jin, Zhe, et al.
Pubblicazione: (2025)
NextMem: Towards Latent Factual Memory for LLM-based Agents
di: Zhang, Zeyu, et al.
Pubblicazione: (2026)
di: Zhang, Zeyu, et al.
Pubblicazione: (2026)
Distilling Transitional Pattern to Large Language Models for Multimodal Session-based Recommendation
di: Su, Jiajie, et al.
Pubblicazione: (2025)
di: Su, Jiajie, et al.
Pubblicazione: (2025)
CARL: Criticality-Aware Agentic Reinforcement Learning
di: Shen, Leyang, et al.
Pubblicazione: (2025)
di: Shen, Leyang, et al.
Pubblicazione: (2025)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
di: Zhang, An, et al.
Pubblicazione: (2024)
di: Zhang, An, et al.
Pubblicazione: (2024)
Reasoning on Time-Series for Financial Technical Analysis
di: Koa, Kelvin J. L., et al.
Pubblicazione: (2025)
di: Koa, Kelvin J. L., et al.
Pubblicazione: (2025)
Post-Training Statistical Calibration for Higher Activation Sparsity
di: Chua, Vui Seng, et al.
Pubblicazione: (2024)
di: Chua, Vui Seng, et al.
Pubblicazione: (2024)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
di: Guo, Shasha, et al.
Pubblicazione: (2024)
di: Guo, Shasha, et al.
Pubblicazione: (2024)
Distillation Enhanced Generative Retrieval
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
On Generative Agents in Recommendation
di: Zhang, An, et al.
Pubblicazione: (2023)
di: Zhang, An, et al.
Pubblicazione: (2023)
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
di: Wang, Lin, et al.
Pubblicazione: (2026)
di: Wang, Lin, et al.
Pubblicazione: (2026)
GOODAT: Towards Test-time Graph Out-of-Distribution Detection
di: Wang, Luzhi, et al.
Pubblicazione: (2024)
di: Wang, Luzhi, et al.
Pubblicazione: (2024)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
di: Zhou, Yang, et al.
Pubblicazione: (2025)
di: Zhou, Yang, et al.
Pubblicazione: (2025)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
di: Yang, Wenjie, et al.
Pubblicazione: (2025)
di: Yang, Wenjie, et al.
Pubblicazione: (2025)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation
di: Zheng, Jingnan, et al.
Pubblicazione: (2024)
di: Zheng, Jingnan, et al.
Pubblicazione: (2024)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
di: Guo, Haiyun, et al.
Pubblicazione: (2025)
di: Guo, Haiyun, et al.
Pubblicazione: (2025)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
di: Liu, Dengcan, et al.
Pubblicazione: (2026)
di: Liu, Dengcan, et al.
Pubblicazione: (2026)
Temporal Relational Reasoning of Large Language Models for Detecting Stock Portfolio Crashes
di: Koa, Kelvin J. L., et al.
Pubblicazione: (2024)
di: Koa, Kelvin J. L., et al.
Pubblicazione: (2024)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
A Distillation-based Future-aware Graph Neural Network for Stock Trend Prediction
di: Liu, Zhipeng, et al.
Pubblicazione: (2025)
di: Liu, Zhipeng, et al.
Pubblicazione: (2025)
Language Representations Can be What Recommenders Need: Findings and Potentials
di: Sheng, Leheng, et al.
Pubblicazione: (2024)
di: Sheng, Leheng, et al.
Pubblicazione: (2024)
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
di: Deng, Yang, et al.
Pubblicazione: (2023)
di: Deng, Yang, et al.
Pubblicazione: (2023)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
di: Chen, Jinyin, et al.
Pubblicazione: (2024)
di: Chen, Jinyin, et al.
Pubblicazione: (2024)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
di: Sharma, Manasi, et al.
Pubblicazione: (2025)
di: Sharma, Manasi, et al.
Pubblicazione: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
di: Li, Gengsheng, et al.
Pubblicazione: (2026) -
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
di: Fang, Junfeng, et al.
Pubblicazione: (2026) -
SOD: Step-wise On-policy Distillation for Small Language Model Agents
di: Zhong, Qiyong, et al.
Pubblicazione: (2026) -
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
di: Sheng, Leheng, et al.
Pubblicazione: (2026) -
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
di: Fang, Junfeng, et al.
Pubblicazione: (2025)