Positive-Unlabeled Reinforcement Learning Distillation for On-Premise Small Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Kou, Zhiqiang, Chen, Junyang, Cai, Xin-Qiang, Xia, Xiaobo, Xie, Ming-Kun, Wu, Dong-Dong, Liu, Biao, Jia, Yuheng, Geng, Xin, Sugiyama, Masashi, Chua, Tat-Seng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Offline Reinforcement Learning with Domain-Unlabeled Data
di: Nishimori, Soichiro, et al.
Pubblicazione: (2024)
di: Nishimori, Soichiro, et al.
Pubblicazione: (2024)
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025)
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025)
Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
di: Wang, Wei, et al.
Pubblicazione: (2025)
di: Wang, Wei, et al.
Pubblicazione: (2025)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2026)
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2026)
Inaccurate Label Distribution Learning with Dependency Noise
di: Kou, Zhiqiang, et al.
Pubblicazione: (2024)
di: Kou, Zhiqiang, et al.
Pubblicazione: (2024)
Label Distribution Learning with Biased Annotations by Learning Multi-Label Representation
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025)
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025)
Principled Multimodal Representation Learning
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Continual Multimodal Contrastive Learning
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Towards Better Performance in Incomplete LDL: Addressing Data Imbalance
di: Kou, Zhiqiang, et al.
Pubblicazione: (2024)
di: Kou, Zhiqiang, et al.
Pubblicazione: (2024)
Are Multimodal Large Language Models Good Annotators for Image Tagging?
di: Xie, Ming-Kun, et al.
Pubblicazione: (2026)
di: Xie, Ming-Kun, et al.
Pubblicazione: (2026)
Towards Modality Generalization: A Benchmark and Prospective Analysis
di: Liu, Xiaohao, et al.
Pubblicazione: (2024)
di: Liu, Xiaohao, et al.
Pubblicazione: (2024)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2025)
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2025)
Reinforcement Learning from Bagged Reward
di: Tang, Yuting, et al.
Pubblicazione: (2024)
di: Tang, Yuting, et al.
Pubblicazione: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
di: Jin, Zhe, et al.
Pubblicazione: (2025)
di: Jin, Zhe, et al.
Pubblicazione: (2025)
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
Learning Robust Diffusion Models from Imprecise Supervision
di: Wu, Dong-Dong, et al.
Pubblicazione: (2025)
di: Wu, Dong-Dong, et al.
Pubblicazione: (2025)
Beyond Simple Sum of Delayed Rewards: Non-Markovian Reward Modeling for Reinforcement Learning
di: Tang, Yuting, et al.
Pubblicazione: (2024)
di: Tang, Yuting, et al.
Pubblicazione: (2024)
Distillation Enhanced Generative Retrieval
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
di: Xie, Ming-Kun, et al.
Pubblicazione: (2025)
di: Xie, Ming-Kun, et al.
Pubblicazione: (2025)
FedHarmony: Harmonizing Heterogeneous Label Correlations in Federated Multi-Label Learning
di: Kou, Zhiqiang, et al.
Pubblicazione: (2026)
di: Kou, Zhiqiang, et al.
Pubblicazione: (2026)
AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
di: Chu, Meng, et al.
Pubblicazione: (2025)
di: Chu, Meng, et al.
Pubblicazione: (2025)
Universal Scene Graph Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
Learning to Ask Critical Questions for Assisting Product Search
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
Trustworthy Federated Label Distribution Learning under Annotation Quality Disparity
di: Wu, Junxiang, et al.
Pubblicazione: (2026)
di: Wu, Junxiang, et al.
Pubblicazione: (2026)
Calibrated Multimodal Representation Learning with Missing Modalities
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
CARL: Criticality-Aware Agentic Reinforcement Learning
di: Shen, Leyang, et al.
Pubblicazione: (2025)
di: Shen, Leyang, et al.
Pubblicazione: (2025)
Multi-Label Knowledge Distillation
di: Yang, Penghui, et al.
Pubblicazione: (2023)
di: Yang, Penghui, et al.
Pubblicazione: (2023)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
Can Class-Priors Help Single-Positive Multi-Label Learning?
di: Liu, Biao, et al.
Pubblicazione: (2023)
di: Liu, Biao, et al.
Pubblicazione: (2023)
Reinforcement Learning with Options and State Representation
di: Ghriss, Ayoub, et al.
Pubblicazione: (2024)
di: Ghriss, Ayoub, et al.
Pubblicazione: (2024)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
di: Shi, Enyi, et al.
Pubblicazione: (2026)
di: Shi, Enyi, et al.
Pubblicazione: (2026)
Positive and Unlabeled Data: Model, Estimation, Inference, and Classification
di: Liu, Siyan, et al.
Pubblicazione: (2024)
di: Liu, Siyan, et al.
Pubblicazione: (2024)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
di: Deng, Yang, et al.
Pubblicazione: (2023)
di: Deng, Yang, et al.
Pubblicazione: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
di: Zhang, An, et al.
Pubblicazione: (2024)
di: Zhang, An, et al.
Pubblicazione: (2024)
Addressing Skewed Heterogeneity via Federated Prototype Rectification with Personalization
di: Guo, Shunxin, et al.
Pubblicazione: (2024)
di: Guo, Shunxin, et al.
Pubblicazione: (2024)
Causal Distillation for Alleviating Performance Heterogeneity in Recommender Systems
di: Zhang, Shengyu, et al.
Pubblicazione: (2024)
di: Zhang, Shengyu, et al.
Pubblicazione: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
di: Ackermann, Johannes, et al.
Pubblicazione: (2024)
di: Ackermann, Johannes, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Offline Reinforcement Learning with Domain-Unlabeled Data
di: Nishimori, Soichiro, et al.
Pubblicazione: (2024) -
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025) -
Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
di: Wang, Wei, et al.
Pubblicazione: (2025) -
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2026) -
Inaccurate Label Distribution Learning with Dependency Noise
di: Kou, Zhiqiang, et al.
Pubblicazione: (2024)