Label-Free Reinforcement Learning via Cross-Model Entropy
Fuente:
arXiv
Guardado en:
| Autores principales: | Gorbett, Matt, Shirazi, Hossein |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
por: Gorbett, Matt, et al.
Publicado: (2024)
por: Gorbett, Matt, et al.
Publicado: (2024)
Cross-Model Disagreement as a Label-Free Correctness Signal
por: Gorbett, Matt, et al.
Publicado: (2026)
por: Gorbett, Matt, et al.
Publicado: (2026)
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
por: Gorbett, Matt, et al.
Publicado: (2023)
por: Gorbett, Matt, et al.
Publicado: (2023)
DISCO-TAB: A Hierarchical Reinforcement Learning Framework for Privacy-Preserving Synthesis of Complex Clinical Data
por: Ilaty, Arshia, et al.
Publicado: (2026)
por: Ilaty, Arshia, et al.
Publicado: (2026)
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
por: Roy, Shuvendu, et al.
Publicado: (2025)
por: Roy, Shuvendu, et al.
Publicado: (2025)
Characterizing Linear Alignment Across Language Models
por: Gorbett, Matt, et al.
Publicado: (2026)
por: Gorbett, Matt, et al.
Publicado: (2026)
Addressing Label Shift in Distributed Learning via Entropy Regularization
por: Wu, Zhiyuan, et al.
Publicado: (2025)
por: Wu, Zhiyuan, et al.
Publicado: (2025)
Entropy-Preserving Reinforcement Learning
por: Petrenko, Aleksei, et al.
Publicado: (2026)
por: Petrenko, Aleksei, et al.
Publicado: (2026)
Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models
por: Yoon, Sangwoong, et al.
Publicado: (2024)
por: Yoon, Sangwoong, et al.
Publicado: (2024)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
por: Jang, Sooyoung, et al.
Publicado: (2021)
por: Jang, Sooyoung, et al.
Publicado: (2021)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
por: Li, Zeqiao, et al.
Publicado: (2026)
por: Li, Zeqiao, et al.
Publicado: (2026)
Learning to Clean: Reinforcement Learning for Noisy Label Correction
por: Heidari, Marzi, et al.
Publicado: (2025)
por: Heidari, Marzi, et al.
Publicado: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
por: Dong, Xiaoyi, et al.
Publicado: (2025)
por: Dong, Xiaoyi, et al.
Publicado: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
por: Bush, Thomas, et al.
Publicado: (2025)
por: Bush, Thomas, et al.
Publicado: (2025)
Towards General-Purpose Model-Free Reinforcement Learning
por: Fujimoto, Scott, et al.
Publicado: (2025)
por: Fujimoto, Scott, et al.
Publicado: (2025)
Smoothie: Label Free Language Model Routing
por: Guha, Neel, et al.
Publicado: (2024)
por: Guha, Neel, et al.
Publicado: (2024)
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
por: Ghanem, Abdelghani, et al.
Publicado: (2026)
por: Ghanem, Abdelghani, et al.
Publicado: (2026)
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis
por: Wang, Yudan, et al.
Publicado: (2024)
por: Wang, Yudan, et al.
Publicado: (2024)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
por: Cui, Ganqu, et al.
Publicado: (2025)
por: Cui, Ganqu, et al.
Publicado: (2025)
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
por: Jin, Renren, et al.
Publicado: (2025)
por: Jin, Renren, et al.
Publicado: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
por: Kim, Dongyoung, et al.
Publicado: (2023)
por: Kim, Dongyoung, et al.
Publicado: (2023)
Reward-Punishment Reinforcement Learning with Maximum Entropy
por: Wang, Jiexin, et al.
Publicado: (2024)
por: Wang, Jiexin, et al.
Publicado: (2024)
A Reinforcement Learning-Based Task Mapping Method to Improve the Reliability of Clustered Manycores
por: Hossein-Khani, Fatemeh, et al.
Publicado: (2024)
por: Hossein-Khani, Fatemeh, et al.
Publicado: (2024)
LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement Learning
por: Ye, Zhuorui, et al.
Publicado: (2024)
por: Ye, Zhuorui, et al.
Publicado: (2024)
SymCircuit: Bayesian Structure Inference for Tractable Probabilistic Circuits via Entropy-Regularized Reinforcement Learning
por: Ju, Y. Sungtaek
Publicado: (2026)
por: Ju, Y. Sungtaek
Publicado: (2026)
Learning Rewards, Not Labels: Adversarial Inverse Reinforcement Learning for Machinery Fault Detection
por: Neupane, Dhiraj, et al.
Publicado: (2026)
por: Neupane, Dhiraj, et al.
Publicado: (2026)
ORAN-GUIDE: RAG-Driven Prompt Learning for LLM-Augmented Reinforcement Learning in O-RAN Network Slicing
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
por: Wang, Shumin, et al.
Publicado: (2026)
por: Wang, Shumin, et al.
Publicado: (2026)
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
por: Ghasemi, Majid, et al.
Publicado: (2024)
por: Ghasemi, Majid, et al.
Publicado: (2024)
Efficient Triple Modular Redundancy for Reliability Enhancement of DNNs Using Explainable AI
por: Soroush, Kimia, et al.
Publicado: (2025)
por: Soroush, Kimia, et al.
Publicado: (2025)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
por: Xie, Sean, et al.
Publicado: (2022)
por: Xie, Sean, et al.
Publicado: (2022)
Reinforcing privacy reasoning in LLMs via normative simulacra from fiction
por: Franchi, Matt, et al.
Publicado: (2026)
por: Franchi, Matt, et al.
Publicado: (2026)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
por: Zhao, Chu, et al.
Publicado: (2026)
por: Zhao, Chu, et al.
Publicado: (2026)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
por: Sanokowski, Sebastian, et al.
Publicado: (2025)
por: Sanokowski, Sebastian, et al.
Publicado: (2025)
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms
por: Aggarwal, Vaneet, et al.
Publicado: (2024)
por: Aggarwal, Vaneet, et al.
Publicado: (2024)
Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions
por: Abate, Alessandro, et al.
Publicado: (2026)
por: Abate, Alessandro, et al.
Publicado: (2026)
Regret-Free Reinforcement Learning for LTL Specifications
por: Majumdar, Rupak, et al.
Publicado: (2024)
por: Majumdar, Rupak, et al.
Publicado: (2024)
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
por: Liu, Wei, et al.
Publicado: (2025)
por: Liu, Wei, et al.
Publicado: (2025)
Source-Free Cross-Domain Continual Learning
por: Furqon, Muhammad Tanzil, et al.
Publicado: (2025)
por: Furqon, Muhammad Tanzil, et al.
Publicado: (2025)
Training-Free Geospatial Place Representation Learning from Large-Scale Point-of-Interest Graph Data
por: Hashemi, Mohammad, et al.
Publicado: (2025)
por: Hashemi, Mohammad, et al.
Publicado: (2025)
Ejemplares similares
-
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
por: Gorbett, Matt, et al.
Publicado: (2024) -
Cross-Model Disagreement as a Label-Free Correctness Signal
por: Gorbett, Matt, et al.
Publicado: (2026) -
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
por: Gorbett, Matt, et al.
Publicado: (2023) -
DISCO-TAB: A Hierarchical Reinforcement Learning Framework for Privacy-Preserving Synthesis of Complex Clinical Data
por: Ilaty, Arshia, et al.
Publicado: (2026) -
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
por: Roy, Shuvendu, et al.
Publicado: (2025)