Guardado en:
| Autores principales: | Petrowski, Michael, Gašić, Milica |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2601.03389 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2024)
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2024)
Dialogue Ontology Relation Extraction via Constrained Chain-of-Thought Decoding
por: Vukovic, Renato, et al.
Publicado: (2024)
por: Vukovic, Renato, et al.
Publicado: (2024)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
por: Zhang, Junkai, et al.
Publicado: (2024)
por: Zhang, Junkai, et al.
Publicado: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
por: Qu, Yuxiao, et al.
Publicado: (2024)
por: Qu, Yuxiao, et al.
Publicado: (2024)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
Latent Introspection: Models Can Detect Prior Concept Injections
por: Pearson-Vogel, Theia, et al.
Publicado: (2026)
por: Pearson-Vogel, Theia, et al.
Publicado: (2026)
Exploration by Random Reward Perturbation
por: Ma, Haozhe, et al.
Publicado: (2025)
por: Ma, Haozhe, et al.
Publicado: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2025)
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2025)
BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
por: Lee, Sangyun, et al.
Publicado: (2025)
por: Lee, Sangyun, et al.
Publicado: (2025)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
por: Kim, Sunghwan, et al.
Publicado: (2025)
por: Kim, Sunghwan, et al.
Publicado: (2025)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
por: Yang, Daniel, et al.
Publicado: (2026)
por: Yang, Daniel, et al.
Publicado: (2026)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
por: Russo, Alessio, et al.
Publicado: (2025)
por: Russo, Alessio, et al.
Publicado: (2025)
Provably Efficient Exploration in Reward Machines with Low Regret
por: Bourel, Hippolyte, et al.
Publicado: (2024)
por: Bourel, Hippolyte, et al.
Publicado: (2024)
Reward Sharpness-Aware Fine-Tuning for Diffusion Models
por: Kim, Kwanyoung, et al.
Publicado: (2026)
por: Kim, Kwanyoung, et al.
Publicado: (2026)
Fairness Aware Reward Optimization
por: Choi, Ching Lam, et al.
Publicado: (2026)
por: Choi, Ching Lam, et al.
Publicado: (2026)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
por: Liang, Yu, et al.
Publicado: (2026)
por: Liang, Yu, et al.
Publicado: (2026)
Zero-Overhead Introspection for Adaptive Test-Time Compute
por: Manvi, Rohin, et al.
Publicado: (2025)
por: Manvi, Rohin, et al.
Publicado: (2025)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
por: Kim, Yoonjeon, et al.
Publicado: (2025)
por: Kim, Yoonjeon, et al.
Publicado: (2025)
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals
por: Liu, Grace, et al.
Publicado: (2024)
por: Liu, Grace, et al.
Publicado: (2024)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
por: Feng, Xiao, et al.
Publicado: (2026)
por: Feng, Xiao, et al.
Publicado: (2026)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
por: Liang, Kaiqu, et al.
Publicado: (2024)
por: Liang, Kaiqu, et al.
Publicado: (2024)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
por: Yang, Diji, et al.
Publicado: (2024)
por: Yang, Diji, et al.
Publicado: (2024)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
por: Cui, Brandon, et al.
Publicado: (2026)
por: Cui, Brandon, et al.
Publicado: (2026)
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
por: He, Qianxi, et al.
Publicado: (2025)
por: He, Qianxi, et al.
Publicado: (2025)
Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards
por: Mohamed, Faisal, et al.
Publicado: (2026)
por: Mohamed, Faisal, et al.
Publicado: (2026)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
por: Li, Mengqi, et al.
Publicado: (2025)
por: Li, Mengqi, et al.
Publicado: (2025)
GRPO is Secretly a Process Reward Model
por: Sullivan, Michael, et al.
Publicado: (2025)
por: Sullivan, Michael, et al.
Publicado: (2025)
SALMON: Self-Alignment with Instructable Reward Models
por: Sun, Zhiqing, et al.
Publicado: (2023)
por: Sun, Zhiqing, et al.
Publicado: (2023)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
por: Jang, Sooyoung, et al.
Publicado: (2021)
por: Jang, Sooyoung, et al.
Publicado: (2021)
Explorations of Self-Repair in Language Models
por: Rushing, Cody, et al.
Publicado: (2024)
por: Rushing, Cody, et al.
Publicado: (2024)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
por: Hu, Xiaomeng, et al.
Publicado: (2025)
por: Hu, Xiaomeng, et al.
Publicado: (2025)
RLSR: Reinforcement Learning from Self Reward
por: Simonds, Toby, et al.
Publicado: (2025)
por: Simonds, Toby, et al.
Publicado: (2025)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
por: Su, Xuerui, et al.
Publicado: (2025)
por: Su, Xuerui, et al.
Publicado: (2025)
ProgAgent:A Continual RL Agent with Progress-Aware Rewards
por: Tan, Jinzhou, et al.
Publicado: (2026)
por: Tan, Jinzhou, et al.
Publicado: (2026)
SemiReward: A General Reward Model for Semi-supervised Learning
por: Li, Siyuan, et al.
Publicado: (2023)
por: Li, Siyuan, et al.
Publicado: (2023)
Self-Generated Critiques Boost Reward Modeling for Language Models
por: Yu, Yue, et al.
Publicado: (2024)
por: Yu, Yue, et al.
Publicado: (2024)
MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward
por: Chen, Kesheng, et al.
Publicado: (2025)
por: Chen, Kesheng, et al.
Publicado: (2025)
Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training
por: Long, Yunbo, et al.
Publicado: (2026)
por: Long, Yunbo, et al.
Publicado: (2026)
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
por: Bhattacharjee, Payel, et al.
Publicado: (2026)
por: Bhattacharjee, Payel, et al.
Publicado: (2026)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
por: Tejaswi, Atula, et al.
Publicado: (2026)
por: Tejaswi, Atula, et al.
Publicado: (2026)
Ejemplares similares
-
Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2024) -
Dialogue Ontology Relation Extraction via Constrained Chain-of-Thought Decoding
por: Vukovic, Renato, et al.
Publicado: (2024) -
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
por: Zhang, Junkai, et al.
Publicado: (2024) -
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
por: Qu, Yuxiao, et al.
Publicado: (2024) -
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)