Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Çağatan, Ömer Veysel, Akgün, Barış |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Failure Modes of Maximum Entropy RLHF
por: Çağatan, Ömer Veysel, et al.
Publicado: (2025)
por: Çağatan, Ömer Veysel, et al.
Publicado: (2025)
Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning
por: Trumpp, Raphael, et al.
Publicado: (2026)
por: Trumpp, Raphael, et al.
Publicado: (2026)
Clipping-Free Policy Optimization for Large Language Models
por: Çağatan, Ömer Veysel, et al.
Publicado: (2026)
por: Çağatan, Ömer Veysel, et al.
Publicado: (2026)
SigCLR: Sigmoid Contrastive Learning of Visual Representations
por: Çağatan, Ömer Veysel
Publicado: (2024)
por: Çağatan, Ömer Veysel
Publicado: (2024)
UNSEE: Unsupervised Non-contrastive Sentence Embeddings
por: Çağatan, Ömer Veysel
Publicado: (2024)
por: Çağatan, Ömer Veysel
Publicado: (2024)
NeoRL: Efficient Exploration for Nonepisodic RL
por: Sukhija, Bhavya, et al.
Publicado: (2024)
por: Sukhija, Bhavya, et al.
Publicado: (2024)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
por: Hu, Zhengding, et al.
Publicado: (2026)
por: Hu, Zhengding, et al.
Publicado: (2026)
Deer-Cora/RL-SSL: RL-SSL
por: Deer-Cora
Publicado: (2026)
por: Deer-Cora
Publicado: (2026)
Task-Induced Representational Invariances Depend on Learning Objective in Deep RL
por: Halvagal, Manu Srinath, et al.
Publicado: (2026)
por: Halvagal, Manu Srinath, et al.
Publicado: (2026)
Improving Transformer World Models for Data-Efficient RL
por: Dedieu, Antoine, et al.
Publicado: (2025)
por: Dedieu, Antoine, et al.
Publicado: (2025)
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
por: Luo, Fan-Ming, et al.
Publicado: (2024)
por: Luo, Fan-Ming, et al.
Publicado: (2024)
Integrating Domain Knowledge for handling Limited Data in Offline RL
por: Gangopadhyay, Briti, et al.
Publicado: (2024)
por: Gangopadhyay, Briti, et al.
Publicado: (2024)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
por: Jackermeier, Mathias, et al.
Publicado: (2024)
por: Jackermeier, Mathias, et al.
Publicado: (2024)
From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
por: Admoni, Sahar, et al.
Publicado: (2025)
por: Admoni, Sahar, et al.
Publicado: (2025)
Efficient RL Training for LLMs with Experience Replay
por: Arnal, Charles, et al.
Publicado: (2026)
por: Arnal, Charles, et al.
Publicado: (2026)
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
por: Zhang, Xikai, et al.
Publicado: (2026)
por: Zhang, Xikai, et al.
Publicado: (2026)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
por: Pavlenko, Kirill, et al.
Publicado: (2026)
por: Pavlenko, Kirill, et al.
Publicado: (2026)
Training RL Agents for Multi-Objective Network Defense Tasks
por: Molina-Markham, Andres, et al.
Publicado: (2025)
por: Molina-Markham, Andres, et al.
Publicado: (2025)
Adversarial Robustness of Discriminative Self-Supervised Learning in Vision
por: Çağatan, Ömer Veysel, et al.
Publicado: (2025)
por: Çağatan, Ömer Veysel, et al.
Publicado: (2025)
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
por: Hung, Yu-Heng, et al.
Publicado: (2025)
por: Hung, Yu-Heng, et al.
Publicado: (2025)
Token-Efficient RL for LLM Reasoning
por: Lee, Alan, et al.
Publicado: (2025)
por: Lee, Alan, et al.
Publicado: (2025)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
por: Bhatia, Abhinav, et al.
Publicado: (2023)
por: Bhatia, Abhinav, et al.
Publicado: (2023)
Uniformly Safe RL with Objective Suppression for Multi-Constraint Safety-Critical Applications
por: Zhou, Zihan, et al.
Publicado: (2024)
por: Zhou, Zihan, et al.
Publicado: (2024)
Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
por: Kim, Sung-Hyun, et al.
Publicado: (2025)
por: Kim, Sung-Hyun, et al.
Publicado: (2025)
Augmenting Offline RL with Unlabeled Data
por: Wang, Zhao, et al.
Publicado: (2024)
por: Wang, Zhao, et al.
Publicado: (2024)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
por: Li, Yuhang, et al.
Publicado: (2026)
por: Li, Yuhang, et al.
Publicado: (2026)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
por: Mark, Max Sobol, et al.
Publicado: (2024)
por: Mark, Max Sobol, et al.
Publicado: (2024)
RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood
por: Zheng, Yifu
Publicado: (2026)
por: Zheng, Yifu
Publicado: (2026)
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
por: Huang, Ruiquan, et al.
Publicado: (2025)
por: Huang, Ruiquan, et al.
Publicado: (2025)
MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster
por: Feng, Laingjun, et al.
Publicado: (2025)
por: Feng, Laingjun, et al.
Publicado: (2025)
Demonstration-Regularized RL
por: Tiapkin, Daniil, et al.
Publicado: (2023)
por: Tiapkin, Daniil, et al.
Publicado: (2023)
RL-Guided Data Selection for Language Model Finetuning
por: Jha, Animesh, et al.
Publicado: (2025)
por: Jha, Animesh, et al.
Publicado: (2025)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
por: Wang, Chen, et al.
Publicado: (2025)
por: Wang, Chen, et al.
Publicado: (2025)
When Are RL Hyperparameters Benign? A Study in Offline Goal-Conditioned RL
por: Töpperwien, Jan Malte, et al.
Publicado: (2026)
por: Töpperwien, Jan Malte, et al.
Publicado: (2026)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
por: Liu, Shaoteng, et al.
Publicado: (2024)
por: Liu, Shaoteng, et al.
Publicado: (2024)
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
por: Cui, Peng, et al.
Publicado: (2026)
por: Cui, Peng, et al.
Publicado: (2026)
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
por: Zhao, Yang, et al.
Publicado: (2025)
por: Zhao, Yang, et al.
Publicado: (2025)
CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
por: Chen, Yutong, et al.
Publicado: (2026)
por: Chen, Yutong, et al.
Publicado: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
por: Xu, Charles, et al.
Publicado: (2026)
por: Xu, Charles, et al.
Publicado: (2026)
BoreaRL: A Multi-Objective Reinforcement Learning Environment for Climate-Adaptive Boreal Forest Management
por: Dsouza, Kevin Bradley, et al.
Publicado: (2025)
por: Dsouza, Kevin Bradley, et al.
Publicado: (2025)
Ejemplares similares
-
Failure Modes of Maximum Entropy RLHF
por: Çağatan, Ömer Veysel, et al.
Publicado: (2025) -
Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning
por: Trumpp, Raphael, et al.
Publicado: (2026) -
Clipping-Free Policy Optimization for Large Language Models
por: Çağatan, Ömer Veysel, et al.
Publicado: (2026) -
SigCLR: Sigmoid Contrastive Learning of Visual Representations
por: Çağatan, Ömer Veysel
Publicado: (2024) -
UNSEE: Unsupervised Non-contrastive Sentence Embeddings
por: Çağatan, Ömer Veysel
Publicado: (2024)