CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Kwon, Soo Min, Sun, Ziteng, Suresh, Ananda Theertha, Jain, Himanshu, Kumar, Sanjiv |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
por: You, Chong, et al.
Publicado: (2025)
por: You, Chong, et al.
Publicado: (2025)
dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models
por: Wan, Zhengyan, et al.
Publicado: (2026)
por: Wan, Zhengyan, et al.
Publicado: (2026)
The importance of feature preprocessing for differentially private linear optimization
por: Sun, Ziteng, et al.
Publicado: (2023)
por: Sun, Ziteng, et al.
Publicado: (2023)
Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
por: Li, Kevin Y., et al.
Publicado: (2026)
por: Li, Kevin Y., et al.
Publicado: (2026)
SpecTr: Fast Speculative Decoding via Optimal Transport
por: Sun, Ziteng, et al.
Publicado: (2023)
por: Sun, Ziteng, et al.
Publicado: (2023)
Subset-Based Instance Optimality in Private Estimation
por: Dick, Travis, et al.
Publicado: (2023)
por: Dick, Travis, et al.
Publicado: (2023)
Knowledge Distillation with Adapted Weight
por: Wu, Sirong, et al.
Publicado: (2025)
por: Wu, Sirong, et al.
Publicado: (2025)
HPM-KD: Hierarchical Progressive Multi-Teacher Framework for Knowledge Distillation and Efficient Model Compression
por: Haase, Gustavo Coelho, et al.
Publicado: (2025)
por: Haase, Gustavo Coelho, et al.
Publicado: (2025)
Longitudinal Risk Prediction in Mammography with Privileged History Distillation
por: Karimian, Banafsheh, et al.
Publicado: (2026)
por: Karimian, Banafsheh, et al.
Publicado: (2026)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
por: Zhou, Renping, et al.
Publicado: (2025)
por: Zhou, Renping, et al.
Publicado: (2025)
Asymptotics of Language Model Alignment
por: Yang, Joy Qiping, et al.
Publicado: (2024)
por: Yang, Joy Qiping, et al.
Publicado: (2024)
A Generative Framework for Causal Estimation via Importance-Weighted Diffusion Distillation
por: Song, Xinran, et al.
Publicado: (2025)
por: Song, Xinran, et al.
Publicado: (2025)
Policy Learning with Observational Data: The Case of Hepatitis C Treatment for HIV/HCV Co-Infected Patients
por: Langevin, Raphaël
Publicado: (2026)
por: Langevin, Raphaël
Publicado: (2026)
Exploring and Improving Drafts in Blockwise Parallel Decoding
por: Kim, Taehyeon, et al.
Publicado: (2024)
por: Kim, Taehyeon, et al.
Publicado: (2024)
CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
por: Sun, Ziteng, et al.
Publicado: (2025)
por: Sun, Ziteng, et al.
Publicado: (2025)
Mapping Methane -- The Impact of Dairy Farm Practices on Emissions Through Satellite Data and Machine Learning
por: Bi, Hanqing, et al.
Publicado: (2024)
por: Bi, Hanqing, et al.
Publicado: (2024)
Co-Evolving Policy Distillation
por: Gu, Naibin, et al.
Publicado: (2026)
por: Gu, Naibin, et al.
Publicado: (2026)
Beyond Beats: A Recipe to Song Popularity? A machine learning approach
por: Sebastian, Niklas, et al.
Publicado: (2024)
por: Sebastian, Niklas, et al.
Publicado: (2024)
Locally-Deployed Chain-of-Thought (CoT) Reasoning Model in Chemical Engineering: Starting from 30 Experimental Data
por: Zhou, Tianhang, et al.
Publicado: (2025)
por: Zhou, Tianhang, et al.
Publicado: (2025)
SynthTree: Co-supervised Local Model Synthesis for Explainable Prediction
por: Kuriabov, Evgenii, et al.
Publicado: (2024)
por: Kuriabov, Evgenii, et al.
Publicado: (2024)
On Robust Hypothesis Testing with respect to the Hellinger Distance
por: Modak, Eeshan, et al.
Publicado: (2025)
por: Modak, Eeshan, et al.
Publicado: (2025)
Bayesian Safe Policy Learning with Chance Constrained Optimization: Application to Military Security Assessment during the Vietnam War
por: Jia, Zeyang, et al.
Publicado: (2023)
por: Jia, Zeyang, et al.
Publicado: (2023)
Coupling without Communication and Drafter-Invariant Speculative Decoding
por: Daliri, Majid, et al.
Publicado: (2024)
por: Daliri, Majid, et al.
Publicado: (2024)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
por: Mundada, Gagan, et al.
Publicado: (2026)
por: Mundada, Gagan, et al.
Publicado: (2026)
Robust Multi-view Co-expression Network Inference
por: Pandeva, Teodora, et al.
Publicado: (2024)
por: Pandeva, Teodora, et al.
Publicado: (2024)
Profit over Proxies: A Scalable Bayesian Decision Framework for Optimizing Multi-Variant Online Experiments
por: Pillai, Srijesh, et al.
Publicado: (2025)
por: Pillai, Srijesh, et al.
Publicado: (2025)
Spatial Heterogeneity in Climate Risk and Human Flourishing: An Exploration with Generative AI
por: Iacus, Stefano Maria, et al.
Publicado: (2026)
por: Iacus, Stefano Maria, et al.
Publicado: (2026)
Mean estimation in the add-remove model of differential privacy
por: Kulesza, Alex, et al.
Publicado: (2023)
por: Kulesza, Alex, et al.
Publicado: (2023)
Expected Diverse Utility (EDU): Diverse Bayesian Optimization of Expensive Computer Simulators
por: Miller, John Joshua, et al.
Publicado: (2024)
por: Miller, John Joshua, et al.
Publicado: (2024)
Efficient Causal Structure Learning via Modular Subgraph Integration
por: Sun, Haixiang, et al.
Publicado: (2026)
por: Sun, Haixiang, et al.
Publicado: (2026)
Multilevel Stochastic Optimization for Imputation in Massive Medical Data Records
por: Li, Wenrui, et al.
Publicado: (2021)
por: Li, Wenrui, et al.
Publicado: (2021)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
por: Deng, Jingcheng, et al.
Publicado: (2026)
por: Deng, Jingcheng, et al.
Publicado: (2026)
Rate of Model Collapse in Recursive Training
por: Suresh, Ananda Theertha, et al.
Publicado: (2024)
por: Suresh, Ananda Theertha, et al.
Publicado: (2024)
Decision Support for Marketplace Policies under Incomplete Evidence: From Replay to Launch Readiness
por: Shekhar, Prashant, et al.
Publicado: (2026)
por: Shekhar, Prashant, et al.
Publicado: (2026)
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
por: Hou, Wenjin, et al.
Publicado: (2026)
por: Hou, Wenjin, et al.
Publicado: (2026)
Vector Optimization with Gaussian Process Bandits
por: Korkmaz, İlter Onat, et al.
Publicado: (2024)
por: Korkmaz, İlter Onat, et al.
Publicado: (2024)
On the Relation Between Autoencoders and Non-negative Matrix Factorization, and Their Application for Mutational Signature Extraction
por: Egendal, Ida, et al.
Publicado: (2024)
por: Egendal, Ida, et al.
Publicado: (2024)
Group-Aware Matrix Estimation and Latent Subspace Recovery
por: Golubovic, Hamza, et al.
Publicado: (2026)
por: Golubovic, Hamza, et al.
Publicado: (2026)
Optimizing Heat Alert Issuance with Reinforcement Learning
por: Considine, Ellen M., et al.
Publicado: (2023)
por: Considine, Ellen M., et al.
Publicado: (2023)
Weather-Related Crash Risk Forecasting: A Deep Learning Approach for Heterogenous Spatiotemporal Data
por: Ogungbire, Abimbola, et al.
Publicado: (2026)
por: Ogungbire, Abimbola, et al.
Publicado: (2026)
Ejemplares similares
-
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
por: You, Chong, et al.
Publicado: (2025) -
dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models
por: Wan, Zhengyan, et al.
Publicado: (2026) -
The importance of feature preprocessing for differentially private linear optimization
por: Sun, Ziteng, et al.
Publicado: (2023) -
Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
por: Li, Kevin Y., et al.
Publicado: (2026) -
SpecTr: Fast Speculative Decoding via Optimal Transport
por: Sun, Ziteng, et al.
Publicado: (2023)