Positive-Unlabeled Reinforcement Learning Distillation for On-Premise Small Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Kou, Zhiqiang, Chen, Junyang, Cai, Xin-Qiang, Xia, Xiaobo, Xie, Ming-Kun, Wu, Dong-Dong, Liu, Biao, Jia, Yuheng, Geng, Xin, Sugiyama, Masashi, Chua, Tat-Seng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Offline Reinforcement Learning with Domain-Unlabeled Data
por: Nishimori, Soichiro, et al.
Publicado: (2024)
por: Nishimori, Soichiro, et al.
Publicado: (2024)
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
por: Kou, Zhiqiang, et al.
Publicado: (2025)
por: Kou, Zhiqiang, et al.
Publicado: (2025)
Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
por: Wang, Wei, et al.
Publicado: (2025)
por: Wang, Wei, et al.
Publicado: (2025)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
por: Cai, Xin-Qiang, et al.
Publicado: (2026)
por: Cai, Xin-Qiang, et al.
Publicado: (2026)
Inaccurate Label Distribution Learning with Dependency Noise
por: Kou, Zhiqiang, et al.
Publicado: (2024)
por: Kou, Zhiqiang, et al.
Publicado: (2024)
Label Distribution Learning with Biased Annotations by Learning Multi-Label Representation
por: Kou, Zhiqiang, et al.
Publicado: (2025)
por: Kou, Zhiqiang, et al.
Publicado: (2025)
Principled Multimodal Representation Learning
por: Liu, Xiaohao, et al.
Publicado: (2025)
por: Liu, Xiaohao, et al.
Publicado: (2025)
Continual Multimodal Contrastive Learning
por: Liu, Xiaohao, et al.
Publicado: (2025)
por: Liu, Xiaohao, et al.
Publicado: (2025)
Towards Better Performance in Incomplete LDL: Addressing Data Imbalance
por: Kou, Zhiqiang, et al.
Publicado: (2024)
por: Kou, Zhiqiang, et al.
Publicado: (2024)
Are Multimodal Large Language Models Good Annotators for Image Tagging?
por: Xie, Ming-Kun, et al.
Publicado: (2026)
por: Xie, Ming-Kun, et al.
Publicado: (2026)
Towards Modality Generalization: A Benchmark and Prospective Analysis
por: Liu, Xiaohao, et al.
Publicado: (2024)
por: Liu, Xiaohao, et al.
Publicado: (2024)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
por: Cai, Xin-Qiang, et al.
Publicado: (2025)
por: Cai, Xin-Qiang, et al.
Publicado: (2025)
Reinforcement Learning from Bagged Reward
por: Tang, Yuting, et al.
Publicado: (2024)
por: Tang, Yuting, et al.
Publicado: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
por: Jin, Zhe, et al.
Publicado: (2025)
por: Jin, Zhe, et al.
Publicado: (2025)
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
por: Zhou, Zhenglin, et al.
Publicado: (2025)
por: Zhou, Zhenglin, et al.
Publicado: (2025)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
por: Zhou, Zhenglin, et al.
Publicado: (2025)
por: Zhou, Zhenglin, et al.
Publicado: (2025)
Learning Robust Diffusion Models from Imprecise Supervision
por: Wu, Dong-Dong, et al.
Publicado: (2025)
por: Wu, Dong-Dong, et al.
Publicado: (2025)
Beyond Simple Sum of Delayed Rewards: Non-Markovian Reward Modeling for Reinforcement Learning
por: Tang, Yuting, et al.
Publicado: (2024)
por: Tang, Yuting, et al.
Publicado: (2024)
Distillation Enhanced Generative Retrieval
por: Li, Yongqi, et al.
Publicado: (2024)
por: Li, Yongqi, et al.
Publicado: (2024)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
por: Xie, Ming-Kun, et al.
Publicado: (2025)
por: Xie, Ming-Kun, et al.
Publicado: (2025)
FedHarmony: Harmonizing Heterogeneous Label Correlations in Federated Multi-Label Learning
por: Kou, Zhiqiang, et al.
Publicado: (2026)
por: Kou, Zhiqiang, et al.
Publicado: (2026)
AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
por: Zhou, Zhenglin, et al.
Publicado: (2025)
por: Zhou, Zhenglin, et al.
Publicado: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
por: Chu, Meng, et al.
Publicado: (2025)
por: Chu, Meng, et al.
Publicado: (2025)
Universal Scene Graph Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
Learning to Ask Critical Questions for Assisting Product Search
por: Li, Zixuan, et al.
Publicado: (2024)
por: Li, Zixuan, et al.
Publicado: (2024)
Trustworthy Federated Label Distribution Learning under Annotation Quality Disparity
por: Wu, Junxiang, et al.
Publicado: (2026)
por: Wu, Junxiang, et al.
Publicado: (2026)
Calibrated Multimodal Representation Learning with Missing Modalities
por: Liu, Xiaohao, et al.
Publicado: (2025)
por: Liu, Xiaohao, et al.
Publicado: (2025)
CARL: Criticality-Aware Agentic Reinforcement Learning
por: Shen, Leyang, et al.
Publicado: (2025)
por: Shen, Leyang, et al.
Publicado: (2025)
Multi-Label Knowledge Distillation
por: Yang, Penghui, et al.
Publicado: (2023)
por: Yang, Penghui, et al.
Publicado: (2023)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
por: Sheng, Leheng, et al.
Publicado: (2026)
por: Sheng, Leheng, et al.
Publicado: (2026)
Can Class-Priors Help Single-Positive Multi-Label Learning?
por: Liu, Biao, et al.
Publicado: (2023)
por: Liu, Biao, et al.
Publicado: (2023)
Reinforcement Learning with Options and State Representation
por: Ghriss, Ayoub, et al.
Publicado: (2024)
por: Ghriss, Ayoub, et al.
Publicado: (2024)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
por: Luo, Run, et al.
Publicado: (2025)
por: Luo, Run, et al.
Publicado: (2025)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
por: Shi, Enyi, et al.
Publicado: (2026)
por: Shi, Enyi, et al.
Publicado: (2026)
Positive and Unlabeled Data: Model, Estimation, Inference, and Classification
por: Liu, Siyan, et al.
Publicado: (2024)
por: Liu, Siyan, et al.
Publicado: (2024)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
por: Deng, Yang, et al.
Publicado: (2023)
por: Deng, Yang, et al.
Publicado: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
por: Zhang, An, et al.
Publicado: (2024)
por: Zhang, An, et al.
Publicado: (2024)
Addressing Skewed Heterogeneity via Federated Prototype Rectification with Personalization
por: Guo, Shunxin, et al.
Publicado: (2024)
por: Guo, Shunxin, et al.
Publicado: (2024)
Causal Distillation for Alleviating Performance Heterogeneity in Recommender Systems
por: Zhang, Shengyu, et al.
Publicado: (2024)
por: Zhang, Shengyu, et al.
Publicado: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
por: Ackermann, Johannes, et al.
Publicado: (2024)
por: Ackermann, Johannes, et al.
Publicado: (2024)
Ejemplares similares
-
Offline Reinforcement Learning with Domain-Unlabeled Data
por: Nishimori, Soichiro, et al.
Publicado: (2024) -
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
por: Kou, Zhiqiang, et al.
Publicado: (2025) -
Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
por: Wang, Wei, et al.
Publicado: (2025) -
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
por: Cai, Xin-Qiang, et al.
Publicado: (2026) -
Inaccurate Label Distribution Learning with Dependency Noise
por: Kou, Zhiqiang, et al.
Publicado: (2024)