Label-Free Reinforcement Learning via Cross-Model Entropy
Fuente:
arXiv
Saved in:
| Main Authors: | Gorbett, Matt, Shirazi, Hossein |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
by: Gorbett, Matt, et al.
Published: (2024)
by: Gorbett, Matt, et al.
Published: (2024)
Cross-Model Disagreement as a Label-Free Correctness Signal
by: Gorbett, Matt, et al.
Published: (2026)
by: Gorbett, Matt, et al.
Published: (2026)
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
by: Gorbett, Matt, et al.
Published: (2023)
by: Gorbett, Matt, et al.
Published: (2023)
DISCO-TAB: A Hierarchical Reinforcement Learning Framework for Privacy-Preserving Synthesis of Complex Clinical Data
by: Ilaty, Arshia, et al.
Published: (2026)
by: Ilaty, Arshia, et al.
Published: (2026)
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
by: Roy, Shuvendu, et al.
Published: (2025)
by: Roy, Shuvendu, et al.
Published: (2025)
Characterizing Linear Alignment Across Language Models
by: Gorbett, Matt, et al.
Published: (2026)
by: Gorbett, Matt, et al.
Published: (2026)
Addressing Label Shift in Distributed Learning via Entropy Regularization
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Entropy-Preserving Reinforcement Learning
by: Petrenko, Aleksei, et al.
Published: (2026)
by: Petrenko, Aleksei, et al.
Published: (2026)
Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models
by: Yoon, Sangwoong, et al.
Published: (2024)
by: Yoon, Sangwoong, et al.
Published: (2024)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
by: Jang, Sooyoung, et al.
Published: (2021)
by: Jang, Sooyoung, et al.
Published: (2021)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
by: Li, Zeqiao, et al.
Published: (2026)
by: Li, Zeqiao, et al.
Published: (2026)
Learning to Clean: Reinforcement Learning for Noisy Label Correction
by: Heidari, Marzi, et al.
Published: (2025)
by: Heidari, Marzi, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
by: Bush, Thomas, et al.
Published: (2025)
by: Bush, Thomas, et al.
Published: (2025)
Towards General-Purpose Model-Free Reinforcement Learning
by: Fujimoto, Scott, et al.
Published: (2025)
by: Fujimoto, Scott, et al.
Published: (2025)
Smoothie: Label Free Language Model Routing
by: Guha, Neel, et al.
Published: (2024)
by: Guha, Neel, et al.
Published: (2024)
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
by: Ghanem, Abdelghani, et al.
Published: (2026)
by: Ghanem, Abdelghani, et al.
Published: (2026)
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis
by: Wang, Yudan, et al.
Published: (2024)
by: Wang, Yudan, et al.
Published: (2024)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
by: Cui, Ganqu, et al.
Published: (2025)
by: Cui, Ganqu, et al.
Published: (2025)
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
by: Jin, Renren, et al.
Published: (2025)
by: Jin, Renren, et al.
Published: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
A Reinforcement Learning-Based Task Mapping Method to Improve the Reliability of Clustered Manycores
by: Hossein-Khani, Fatemeh, et al.
Published: (2024)
by: Hossein-Khani, Fatemeh, et al.
Published: (2024)
LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement Learning
by: Ye, Zhuorui, et al.
Published: (2024)
by: Ye, Zhuorui, et al.
Published: (2024)
SymCircuit: Bayesian Structure Inference for Tractable Probabilistic Circuits via Entropy-Regularized Reinforcement Learning
by: Ju, Y. Sungtaek
Published: (2026)
by: Ju, Y. Sungtaek
Published: (2026)
Learning Rewards, Not Labels: Adversarial Inverse Reinforcement Learning for Machinery Fault Detection
by: Neupane, Dhiraj, et al.
Published: (2026)
by: Neupane, Dhiraj, et al.
Published: (2026)
ORAN-GUIDE: RAG-Driven Prompt Learning for LLM-Augmented Reinforcement Learning in O-RAN Network Slicing
by: Lotfi, Fatemeh, et al.
Published: (2025)
by: Lotfi, Fatemeh, et al.
Published: (2025)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
by: Wang, Shumin, et al.
Published: (2026)
by: Wang, Shumin, et al.
Published: (2026)
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
by: Ghasemi, Majid, et al.
Published: (2024)
by: Ghasemi, Majid, et al.
Published: (2024)
Efficient Triple Modular Redundancy for Reliability Enhancement of DNNs Using Explainable AI
by: Soroush, Kimia, et al.
Published: (2025)
by: Soroush, Kimia, et al.
Published: (2025)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
by: Xie, Sean, et al.
Published: (2022)
by: Xie, Sean, et al.
Published: (2022)
Reinforcing privacy reasoning in LLMs via normative simulacra from fiction
by: Franchi, Matt, et al.
Published: (2026)
by: Franchi, Matt, et al.
Published: (2026)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
by: Zhao, Chu, et al.
Published: (2026)
by: Zhao, Chu, et al.
Published: (2026)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
by: Sanokowski, Sebastian, et al.
Published: (2025)
by: Sanokowski, Sebastian, et al.
Published: (2025)
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms
by: Aggarwal, Vaneet, et al.
Published: (2024)
by: Aggarwal, Vaneet, et al.
Published: (2024)
Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions
by: Abate, Alessandro, et al.
Published: (2026)
by: Abate, Alessandro, et al.
Published: (2026)
Regret-Free Reinforcement Learning for LTL Specifications
by: Majumdar, Rupak, et al.
Published: (2024)
by: Majumdar, Rupak, et al.
Published: (2024)
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Source-Free Cross-Domain Continual Learning
by: Furqon, Muhammad Tanzil, et al.
Published: (2025)
by: Furqon, Muhammad Tanzil, et al.
Published: (2025)
Training-Free Geospatial Place Representation Learning from Large-Scale Point-of-Interest Graph Data
by: Hashemi, Mohammad, et al.
Published: (2025)
by: Hashemi, Mohammad, et al.
Published: (2025)
Similar Items
-
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
by: Gorbett, Matt, et al.
Published: (2024) -
Cross-Model Disagreement as a Label-Free Correctness Signal
by: Gorbett, Matt, et al.
Published: (2026) -
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
by: Gorbett, Matt, et al.
Published: (2023) -
DISCO-TAB: A Hierarchical Reinforcement Learning Framework for Privacy-Preserving Synthesis of Complex Clinical Data
by: Ilaty, Arshia, et al.
Published: (2026) -
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
by: Roy, Shuvendu, et al.
Published: (2025)