TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoon, Eunseop, Yoon, Hee Suk, Eom, SooHwan, Han, Gunsoo, Nam, Daniel Wontae, Jo, Daejin, On, Kyoung-Woon, Hasegawa-Johnson, Mark A., Kim, Sungwoong, Yoo, Chang D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
Hexa: Self-Improving for Knowledge-Grounded Dialogue System
von: Jo, Daejin, et al.
Veröffentlicht: (2023)
von: Jo, Daejin, et al.
Veröffentlicht: (2023)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
von: Hong, Ji Woo, et al.
Veröffentlicht: (2026)
von: Hong, Ji Woo, et al.
Veröffentlicht: (2026)
Binary Classifier Optimization for Large Language Model Alignment
von: Jung, Seungjae, et al.
Veröffentlicht: (2024)
von: Jung, Seungjae, et al.
Veröffentlicht: (2024)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
von: Yoon, Eunseop, et al.
Veröffentlicht: (2024)
von: Yoon, Eunseop, et al.
Veröffentlicht: (2024)
SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation
von: Ryu, Hyun, et al.
Veröffentlicht: (2023)
von: Ryu, Hyun, et al.
Veröffentlicht: (2023)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2024)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2024)
HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue
von: Yoon, Sunjae, et al.
Veröffentlicht: (2023)
von: Yoon, Sunjae, et al.
Veröffentlicht: (2023)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
von: Yoon, Sunjae, et al.
Veröffentlicht: (2022)
von: Yoon, Sunjae, et al.
Veröffentlicht: (2022)
SGPO: Self-Generated Preference Optimization based on Self-Improver
von: Lee, Hyeonji, et al.
Veröffentlicht: (2025)
von: Lee, Hyeonji, et al.
Veröffentlicht: (2025)
ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2023)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2023)
Zero-Shot Dual-Path Integration Framework for Open-Vocabulary 3D Instance Segmentation
von: Ton, Tri, et al.
Veröffentlicht: (2024)
von: Ton, Tri, et al.
Veröffentlicht: (2024)
BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2024)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
Programmable spectral shaping to improve the measurement precision of frequency comb mode-resolved spectral interferometric ranging
von: Jang, Yoon-Soo, et al.
Veröffentlicht: (2023)
von: Jang, Yoon-Soo, et al.
Veröffentlicht: (2023)
Highly Ordered Mesoporous Polymer‐Supported Phosphine as the Ligand for Organometallic Reaction: Suzuki‐Miyaura Cross‐Coupling of Aryl Chlorides at Room Temperature
von: Hwang Suk Kim, et al.
Veröffentlicht: (2024)
von: Hwang Suk Kim, et al.
Veröffentlicht: (2024)
Term Structure and Risk Premiums of Commodity Futures With Linear Regressions
von: Daejin Kim
Veröffentlicht: (2024)
von: Daejin Kim
Veröffentlicht: (2024)
Validity of black hole complementarity in an accelerating Schwarzschild black hole
von: Kim, Wontae, et al.
Veröffentlicht: (2024)
von: Kim, Wontae, et al.
Veröffentlicht: (2024)
Temperatures of AdS$_2$ black holes and holography revisited
von: Kim, Wontae, et al.
Veröffentlicht: (2023)
von: Kim, Wontae, et al.
Veröffentlicht: (2023)
Gravitational constant as a conserved charge in black hole thermodynamics
von: Kim, Wontae, et al.
Veröffentlicht: (2025)
von: Kim, Wontae, et al.
Veröffentlicht: (2025)
Investigation of black hole complementarity in AdS$_2$ black holes
von: Kim, Wontae, et al.
Veröffentlicht: (2023)
von: Kim, Wontae, et al.
Veröffentlicht: (2023)
The study on the multiplicity dependence of ridge behavior in $pp$ collisions at $\sqrt{s}=13$ TeV at the LHC
von: Yoon, Jeongseok, et al.
Veröffentlicht: (2023)
von: Yoon, Jeongseok, et al.
Veröffentlicht: (2023)
Approaching the quantum-limited precision in frequency-comb-based spectral interferometry for length measurements
von: Jang, Yoon-Soo, et al.
Veröffentlicht: (2025)
von: Jang, Yoon-Soo, et al.
Veröffentlicht: (2025)
Approaching the Quantum‐Limited Precision in Frequency‐Comb‐Based Spectral Interferometric Ranging
von: Yoon‐Soo Jang, et al.
Veröffentlicht: (2025)
von: Yoon‐Soo Jang, et al.
Veröffentlicht: (2025)
Approaching the Quantum‐Limited Precision in Frequency‐Comb‐Based Spectral Interferometric Ranging (Laser Photonics Rev. 19(11)/2025)
von: Yoon‐Soo Jang, et al.
Veröffentlicht: (2025)
von: Yoon‐Soo Jang, et al.
Veröffentlicht: (2025)
Versatile and Fast Location-Based Private Information Retrieval with Fully Homomorphic Encryption over the Torus
von: Yoo, Joon Soo, et al.
Veröffentlicht: (2025)
von: Yoo, Joon Soo, et al.
Veröffentlicht: (2025)
On the Improvement of the "Copyright Law" of Korea for Library Services for Persons with Disabilities
von: Yoon, Hee-Yoon, et al.
Veröffentlicht: (2013)
von: Yoon, Hee-Yoon, et al.
Veröffentlicht: (2013)
Production of all-female diploid and triploid far eastern catfish, Silurus asotus (Linnaeus) : survival and growth performance / Yoon Kwon Nam
von: Kwon Nam, Yoon
Veröffentlicht: (2001)
von: Kwon Nam, Yoon
Veröffentlicht: (2001)
Bidirectional Biometric Authentication Using Transciphering and (T)FHE
von: Yoo, Joon Soo, et al.
Veröffentlicht: (2025)
von: Yoo, Joon Soo, et al.
Veröffentlicht: (2025)
Off–On fluorescent benzothiazole‐fused coumarin for sensitive detection of nitroreductases and hydrogen sulfide
von: Song Yi Yoo, et al.
Veröffentlicht: (2024)
von: Song Yi Yoo, et al.
Veröffentlicht: (2024)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
von: Yune, Sungwoong, et al.
Veröffentlicht: (2026)
von: Yune, Sungwoong, et al.
Veröffentlicht: (2026)
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022)
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022)
Systemic Therapy for Advanced Biliary Tract Cancers in 2026: Current Standard of Care and Emerging Therapeutic Strategies
von: Hyunseok Yoon, et al.
Veröffentlicht: (2025)
von: Hyunseok Yoon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
von: Eom, SooHwan, et al.
Veröffentlicht: (2024) -
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025) -
Hexa: Self-Improving for Knowledge-Grounded Dialogue System
von: Jo, Daejin, et al.
Veröffentlicht: (2023) -
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025) -
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)