Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | de Langis, Karin, Koo, Ryan, Kang, Dongyeop |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
di: de Langis, Karin, et al.
Pubblicazione: (2025)
di: de Langis, Karin, et al.
Pubblicazione: (2025)
Effects of Varying LLM Access on Essay Writing Behavior
di: Christenson, Julia, et al.
Pubblicazione: (2026)
di: Christenson, Julia, et al.
Pubblicazione: (2026)
Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?
di: de Langis, Karin, et al.
Pubblicazione: (2025)
di: de Langis, Karin, et al.
Pubblicazione: (2025)
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs
di: de Langis, Karin, et al.
Pubblicazione: (2025)
di: de Langis, Karin, et al.
Pubblicazione: (2025)
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
di: de Langis, Karin, et al.
Pubblicazione: (2025)
di: de Langis, Karin, et al.
Pubblicazione: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
di: Das, Debarati, et al.
Pubblicazione: (2024)
di: Das, Debarati, et al.
Pubblicazione: (2024)
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
di: Kim, Zae Myung, et al.
Pubblicazione: (2025)
di: Kim, Zae Myung, et al.
Pubblicazione: (2025)
Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
di: Yang, Ruixin, et al.
Pubblicazione: (2024)
di: Yang, Ruixin, et al.
Pubblicazione: (2024)
Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
di: Lu, Yining, et al.
Pubblicazione: (2025)
di: Lu, Yining, et al.
Pubblicazione: (2025)
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
di: Min, Do June, et al.
Pubblicazione: (2024)
di: Min, Do June, et al.
Pubblicazione: (2024)
MPCODER: Multi-user Personalized Code Generator with Explicit and Implicit Style Representation Learning
di: Dai, Zhenlong, et al.
Pubblicazione: (2024)
di: Dai, Zhenlong, et al.
Pubblicazione: (2024)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
di: Koo, Ryan, et al.
Pubblicazione: (2023)
di: Koo, Ryan, et al.
Pubblicazione: (2023)
Shallow Synthesis of Knowledge in GPT-Generated Texts: A Case Study in Automatic Related Work Composition
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2024)
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2024)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
di: Zhai, Skylar, et al.
Pubblicazione: (2026)
di: Zhai, Skylar, et al.
Pubblicazione: (2026)
Scaling Unverifiable Rewards: A Case Study on Visual Insights
di: Gan, Shuyu, et al.
Pubblicazione: (2025)
di: Gan, Shuyu, et al.
Pubblicazione: (2025)
Show and Tell: Prompt Strategies for Style Control in Multi-Turn LLM Code Generation
di: Bohr, Jeremiah
Pubblicazione: (2025)
di: Bohr, Jeremiah
Pubblicazione: (2025)
A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting
di: Gan, Shuyu, et al.
Pubblicazione: (2025)
di: Gan, Shuyu, et al.
Pubblicazione: (2025)
StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing
di: Cong, Gaoxiang, et al.
Pubblicazione: (2024)
di: Cong, Gaoxiang, et al.
Pubblicazione: (2024)
How Far Can We Extract Diverse Perspectives from Large Language Models?
di: Hayati, Shirley Anugrah, et al.
Pubblicazione: (2023)
di: Hayati, Shirley Anugrah, et al.
Pubblicazione: (2023)
LawFlow: Collecting and Simulating Lawyers' Thought Processes on Business Formation Case Studies
di: Das, Debarati, et al.
Pubblicazione: (2025)
di: Das, Debarati, et al.
Pubblicazione: (2025)
Style Transfer with Multi-iteration Preference Optimization
di: Liu, Shuai, et al.
Pubblicazione: (2024)
di: Liu, Shuai, et al.
Pubblicazione: (2024)
Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation
di: Park, Jong Inn, et al.
Pubblicazione: (2025)
di: Park, Jong Inn, et al.
Pubblicazione: (2025)
BBScore: A Brownian Bridge Based Metric for Assessing Text Coherence
di: Sheng, Zhecheng, et al.
Pubblicazione: (2023)
di: Sheng, Zhecheng, et al.
Pubblicazione: (2023)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation
di: Shen, Jiaming, et al.
Pubblicazione: (2024)
di: Shen, Jiaming, et al.
Pubblicazione: (2024)
Threads of Subtlety: Detecting Machine-Generated Texts Through Discourse Motifs
di: Kim, Zae Myung, et al.
Pubblicazione: (2024)
di: Kim, Zae Myung, et al.
Pubblicazione: (2024)
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
di: Li, Xiaomin, et al.
Pubblicazione: (2025)
di: Li, Xiaomin, et al.
Pubblicazione: (2025)
Tailoring Self-Rationalizers with Multi-Reward Distillation
di: Ramnath, Sahana, et al.
Pubblicazione: (2023)
di: Ramnath, Sahana, et al.
Pubblicazione: (2023)
Which Modality should I use -- Text, Motif, or Image? : Understanding Graphs with Large Language Models
di: Das, Debarati, et al.
Pubblicazione: (2023)
di: Das, Debarati, et al.
Pubblicazione: (2023)
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2025)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2025)
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
di: Cheshmi, Seyyed Saeid, et al.
Pubblicazione: (2026)
di: Cheshmi, Seyyed Saeid, et al.
Pubblicazione: (2026)
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
di: Liu, Yantao, et al.
Pubblicazione: (2024)
di: Liu, Yantao, et al.
Pubblicazione: (2024)
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
di: Deng, Haikang, et al.
Pubblicazione: (2023)
di: Deng, Haikang, et al.
Pubblicazione: (2023)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
di: Lee, Minhwa, et al.
Pubblicazione: (2024)
di: Lee, Minhwa, et al.
Pubblicazione: (2024)
Beyond Single-Reward: Multi-Pair, Multi-Perspective Preference Optimization for Machine Translation
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
di: Lin, Yu-Xiang, et al.
Pubblicazione: (2025)
di: Lin, Yu-Xiang, et al.
Pubblicazione: (2025)
Weighted-Reward Preference Optimization for Implicit Model Fusion
di: Yang, Ziyi, et al.
Pubblicazione: (2024)
di: Yang, Ziyi, et al.
Pubblicazione: (2024)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
di: Parkar, Ritik Sachin, et al.
Pubblicazione: (2024)
di: Parkar, Ritik Sachin, et al.
Pubblicazione: (2024)
Structure Liberates: How Constrained Sensemaking Produces More Novel Research Output
di: Mooney, James, et al.
Pubblicazione: (2026)
di: Mooney, James, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
di: de Langis, Karin, et al.
Pubblicazione: (2025) -
Effects of Varying LLM Access on Essay Writing Behavior
di: Christenson, Julia, et al.
Pubblicazione: (2026) -
Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?
di: de Langis, Karin, et al.
Pubblicazione: (2025) -
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs
di: de Langis, Karin, et al.
Pubblicazione: (2025) -
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
di: de Langis, Karin, et al.
Pubblicazione: (2025)