Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
Fuente:
arXiv
Salvato in:
| Autori principali: | Choi, Yunho, Lim, Jongwon, Ahn, Woojin, Oh, Minjae, Shim, Jeonghoon, Jo, Yohan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning
di: Oh, Minjae, et al.
Pubblicazione: (2025)
di: Oh, Minjae, et al.
Pubblicazione: (2025)
KL for a KL: On-Policy Distillation with Control Variate Baseline
di: Oh, Minjae, et al.
Pubblicazione: (2026)
di: Oh, Minjae, et al.
Pubblicazione: (2026)
ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
di: Shim, Jeonghoon, et al.
Pubblicazione: (2025)
di: Shim, Jeonghoon, et al.
Pubblicazione: (2025)
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
di: Park, Sejun, et al.
Pubblicazione: (2026)
di: Park, Sejun, et al.
Pubblicazione: (2026)
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
di: Chung, Woojin, et al.
Pubblicazione: (2025)
di: Chung, Woojin, et al.
Pubblicazione: (2025)
Non-Collaborative User Simulators for Tool Agents
di: Shim, Jeonghoon, et al.
Pubblicazione: (2025)
di: Shim, Jeonghoon, et al.
Pubblicazione: (2025)
ThinkBrake: Efficient Reasoning via Log-Probability Margin Guided Decoding
di: Song, Sangjun, et al.
Pubblicazione: (2025)
di: Song, Sangjun, et al.
Pubblicazione: (2025)
Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance
di: Chung, Woojin, et al.
Pubblicazione: (2026)
di: Chung, Woojin, et al.
Pubblicazione: (2026)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
di: Yu, Dian, et al.
Pubblicazione: (2025)
di: Yu, Dian, et al.
Pubblicazione: (2025)
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
di: Han, Jongwook, et al.
Pubblicazione: (2025)
di: Han, Jongwook, et al.
Pubblicazione: (2025)
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
di: Kim, Woojin, et al.
Pubblicazione: (2026)
di: Kim, Woojin, et al.
Pubblicazione: (2026)
BYOL: Bring Your Own Language Into LLMs
di: Zamir, Syed Waqas, et al.
Pubblicazione: (2026)
di: Zamir, Syed Waqas, et al.
Pubblicazione: (2026)
SUIT: Knowledge Editing with Subspace-Aware Key-Value Mappings
di: Park, Haewon, et al.
Pubblicazione: (2025)
di: Park, Haewon, et al.
Pubblicazione: (2025)
Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline
di: Lee, Tony, et al.
Pubblicazione: (2026)
di: Lee, Tony, et al.
Pubblicazione: (2026)
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
di: Kong, Injin, et al.
Pubblicazione: (2026)
di: Kong, Injin, et al.
Pubblicazione: (2026)
Context-Robust Knowledge Editing for Language Models
di: Park, Haewon, et al.
Pubblicazione: (2025)
di: Park, Haewon, et al.
Pubblicazione: (2025)
Dialogue Systems for Emotional Support via Value Reinforcement
di: Kim, Juhee, et al.
Pubblicazione: (2025)
di: Kim, Juhee, et al.
Pubblicazione: (2025)
Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning
di: Baheri, Ali, et al.
Pubblicazione: (2025)
di: Baheri, Ali, et al.
Pubblicazione: (2025)
Bootstrap Your Own Context Length
di: Wang, Liang, et al.
Pubblicazione: (2024)
di: Wang, Liang, et al.
Pubblicazione: (2024)
Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples
di: Pyun, Haesung, et al.
Pubblicazione: (2025)
di: Pyun, Haesung, et al.
Pubblicazione: (2025)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
di: Hong, Joey, et al.
Pubblicazione: (2025)
di: Hong, Joey, et al.
Pubblicazione: (2025)
Gender Bias in LLM-generated Interview Responses
di: Kong, Haein, et al.
Pubblicazione: (2024)
di: Kong, Haein, et al.
Pubblicazione: (2024)
Contraction Actor-Critic: Contraction Metric-Guided Reinforcement Learning for Robust Path Tracking
di: Cho, Minjae, et al.
Pubblicazione: (2025)
di: Cho, Minjae, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
di: Chae, Kyubyung, et al.
Pubblicazione: (2024)
di: Chae, Kyubyung, et al.
Pubblicazione: (2024)
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
di: Cho, Ye-eun, et al.
Pubblicazione: (2025)
di: Cho, Ye-eun, et al.
Pubblicazione: (2025)
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
di: Yu, Longhui, et al.
Pubblicazione: (2023)
di: Yu, Longhui, et al.
Pubblicazione: (2023)
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
di: Lim, Sungjib, et al.
Pubblicazione: (2025)
di: Lim, Sungjib, et al.
Pubblicazione: (2025)
Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
di: Yoo, YoungJoon, et al.
Pubblicazione: (2023)
di: Yoo, YoungJoon, et al.
Pubblicazione: (2023)
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
di: Kong, Injin, et al.
Pubblicazione: (2026)
di: Kong, Injin, et al.
Pubblicazione: (2026)
Feeding LLM Annotations to BERT Classifiers at Your Own Risk
di: Lu, Yucheng, et al.
Pubblicazione: (2025)
di: Lu, Yucheng, et al.
Pubblicazione: (2025)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
di: Lee, Jonggeun, et al.
Pubblicazione: (2026)
di: Lee, Jonggeun, et al.
Pubblicazione: (2026)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
di: Seo, Jean, et al.
Pubblicazione: (2024)
di: Seo, Jean, et al.
Pubblicazione: (2024)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
di: Choi, Sangun, et al.
Pubblicazione: (2025)
di: Choi, Sangun, et al.
Pubblicazione: (2025)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
di: Park, Yoonah, et al.
Pubblicazione: (2025)
di: Park, Yoonah, et al.
Pubblicazione: (2025)
Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
di: Heo, Chaewon, et al.
Pubblicazione: (2026)
di: Heo, Chaewon, et al.
Pubblicazione: (2026)
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
di: Chen, Zizhe, et al.
Pubblicazione: (2026)
di: Chen, Zizhe, et al.
Pubblicazione: (2026)
Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
di: Ghasemabadi, Amirhosein, et al.
Pubblicazione: (2025)
di: Ghasemabadi, Amirhosein, et al.
Pubblicazione: (2025)
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
di: Choi, Jaepill, et al.
Pubblicazione: (2024)
di: Choi, Jaepill, et al.
Pubblicazione: (2024)
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
di: Wang, Mingyang, et al.
Pubblicazione: (2025)
di: Wang, Mingyang, et al.
Pubblicazione: (2025)
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
di: Lee, Jonggeun, et al.
Pubblicazione: (2025)
di: Lee, Jonggeun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning
di: Oh, Minjae, et al.
Pubblicazione: (2025) -
KL for a KL: On-Policy Distillation with Control Variate Baseline
di: Oh, Minjae, et al.
Pubblicazione: (2026) -
ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
di: Shim, Jeonghoon, et al.
Pubblicazione: (2025) -
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
di: Park, Sejun, et al.
Pubblicazione: (2026) -
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
di: Chung, Woojin, et al.
Pubblicazione: (2025)