Reinforcement Learning with Token-level Feedback for Controllable Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Wendi, Wei, Wei, Xu, Kaihe, Xie, Wenfeng, Chen, Dangyang, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
von: Fan, Shixuan, et al.
Veröffentlicht: (2024)
von: Fan, Shixuan, et al.
Veröffentlicht: (2024)
Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
Enhancing Low-Resource Relation Representations through Multi-View Decoupling
von: Fan, Chenghao, et al.
Veröffentlicht: (2023)
von: Fan, Chenghao, et al.
Veröffentlicht: (2023)
Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph
von: Huang, Rikui, et al.
Veröffentlicht: (2024)
von: Huang, Rikui, et al.
Veröffentlicht: (2024)
Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
von: Ding, Zhuojun, et al.
Veröffentlicht: (2024)
von: Ding, Zhuojun, et al.
Veröffentlicht: (2024)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
Personalized Topic Selection Model for Topic-Grounded Dialogue
von: Fan, Shixuan, et al.
Veröffentlicht: (2024)
von: Fan, Shixuan, et al.
Veröffentlicht: (2024)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
von: Fan, Chenghao, et al.
Veröffentlicht: (2024)
von: Fan, Chenghao, et al.
Veröffentlicht: (2024)
SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL
von: Hua, Harper, et al.
Veröffentlicht: (2026)
von: Hua, Harper, et al.
Veröffentlicht: (2026)
Enhancing the Traditional Chinese Medicine Capabilities of Large Language Model through Reinforcement Learning from AI Feedback
von: Yu, Song, et al.
Veröffentlicht: (2024)
von: Yu, Song, et al.
Veröffentlicht: (2024)
Text2Grad: Reinforcement Learning from Natural Language Feedback
von: Wang, Hanyang, et al.
Veröffentlicht: (2025)
von: Wang, Hanyang, et al.
Veröffentlicht: (2025)
Dr Genre: Reinforcement Learning from Decoupled LLM Feedback for Generic Text Rewriting
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
von: Shen, Wei, et al.
Veröffentlicht: (2024)
von: Shen, Wei, et al.
Veröffentlicht: (2024)
Intelligent Tutor: Leveraging ChatGPT and Microsoft Copilot Studio to Deliver a Generative AI Student Support and Feedback System within Teams
von: Chen, Wei-Yu
Veröffentlicht: (2024)
von: Chen, Wei-Yu
Veröffentlicht: (2024)
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
von: Li, Jiahui, et al.
Veröffentlicht: (2024)
von: Li, Jiahui, et al.
Veröffentlicht: (2024)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
von: Ren, Houxing, et al.
Veröffentlicht: (2024)
von: Ren, Houxing, et al.
Veröffentlicht: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
von: Huang, Lei, et al.
Veröffentlicht: (2026)
von: Huang, Lei, et al.
Veröffentlicht: (2026)
Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
Parameter Efficient Reinforcement Learning from Human Feedback
von: Sidahmed, Hakim, et al.
Veröffentlicht: (2024)
von: Sidahmed, Hakim, et al.
Veröffentlicht: (2024)
Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
von: Tao, Zhen, et al.
Veröffentlicht: (2024)
von: Tao, Zhen, et al.
Veröffentlicht: (2024)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
Process Reward Model with Q-Value Rankings
von: Li, Wendi, et al.
Veröffentlicht: (2024)
von: Li, Wendi, et al.
Veröffentlicht: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
von: Tian, Wei, et al.
Veröffentlicht: (2026)
von: Tian, Wei, et al.
Veröffentlicht: (2026)
Explainability-Based Token Replacement on LLM-Generated Text
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Backtracking Feedback
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
Token Masking Improves Transformer-Based Text Classification
von: Xu, Xianglong, et al.
Veröffentlicht: (2025)
von: Xu, Xianglong, et al.
Veröffentlicht: (2025)
Alternatives To Next Token Prediction In Text Generation -- A Survey
von: Wyatt, Charlie, et al.
Veröffentlicht: (2025)
von: Wyatt, Charlie, et al.
Veröffentlicht: (2025)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
von: Liu, Shiqi, et al.
Veröffentlicht: (2026)
von: Liu, Shiqi, et al.
Veröffentlicht: (2026)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025)
von: Li, Wendi, et al.
Veröffentlicht: (2025)
MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text
von: Li, Chenjun, et al.
Veröffentlicht: (2026)
von: Li, Chenjun, et al.
Veröffentlicht: (2026)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
Factuality on Demand: Controlling the Factuality-Informativeness Trade-off in Text Generation
von: Gong, Ziwei, et al.
Veröffentlicht: (2026)
von: Gong, Ziwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
von: Fan, Shixuan, et al.
Veröffentlicht: (2024) -
Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024) -
Enhancing Low-Resource Relation Representations through Multi-View Decoupling
von: Fan, Chenghao, et al.
Veröffentlicht: (2023) -
Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph
von: Huang, Rikui, et al.
Veröffentlicht: (2024) -
Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
von: Ding, Zhuojun, et al.
Veröffentlicht: (2024)