Users as Annotators: LLM Preference Learning from Comparison Mode
Fuente:
arXiv
Guardado en:
| Autores principales: | Cai, Zhongze, Li, Xiaocheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Goal Alignment in LLM-Based User Simulators for Conversational AI
por: Mehri, Shuhaib, et al.
Publicado: (2025)
por: Mehri, Shuhaib, et al.
Publicado: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024)
por: Kim, Dongyoung, et al.
Publicado: (2024)
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
por: Wang, Yifan, et al.
Publicado: (2025)
por: Wang, Yifan, et al.
Publicado: (2025)
Arithmetic Reasoning with LLM: Prolog Generation & Permutation
por: Yang, Xiaocheng, et al.
Publicado: (2024)
por: Yang, Xiaocheng, et al.
Publicado: (2024)
Aligning LLM Agents by Learning Latent Preference from User Edits
por: Gao, Ge, et al.
Publicado: (2024)
por: Gao, Ge, et al.
Publicado: (2024)
From 1,000,000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment
por: Li, Jia-Nan, et al.
Publicado: (2025)
por: Li, Jia-Nan, et al.
Publicado: (2025)
Dissecting Human and LLM Preferences
por: Li, Junlong, et al.
Publicado: (2024)
por: Li, Junlong, et al.
Publicado: (2024)
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification
por: Liu, Shang, et al.
Publicado: (2024)
por: Liu, Shang, et al.
Publicado: (2024)
LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
por: Yuan, Liangqi, et al.
Publicado: (2025)
por: Yuan, Liangqi, et al.
Publicado: (2025)
Learning from Response not Preference: A Stackelberg Approach for LLM Detoxification using Non-parallel Data
por: Xie, Xinhong, et al.
Publicado: (2024)
por: Xie, Xinhong, et al.
Publicado: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
por: Kim, Kon Woo, et al.
Publicado: (2025)
por: Kim, Kon Woo, et al.
Publicado: (2025)
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
por: Feng, Duanyu, et al.
Publicado: (2024)
por: Feng, Duanyu, et al.
Publicado: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
por: Li, Ziang, et al.
Publicado: (2025)
por: Li, Ziang, et al.
Publicado: (2025)
Enhancing LLM-Based Data Annotation with Error Decomposition
por: Xu, Zhen, et al.
Publicado: (2026)
por: Xu, Zhen, et al.
Publicado: (2026)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
por: Ni, Jingwei, et al.
Publicado: (2024)
por: Ni, Jingwei, et al.
Publicado: (2024)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
por: Zheng, Mingqian, et al.
Publicado: (2025)
por: Zheng, Mingqian, et al.
Publicado: (2025)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
por: Jung, Jeesu, et al.
Publicado: (2025)
por: Jung, Jeesu, et al.
Publicado: (2025)
RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step
por: Luo, Xiaocheng, et al.
Publicado: (2026)
por: Luo, Xiaocheng, et al.
Publicado: (2026)
Emergent Convergence in Multi-Agent LLM Annotation
por: Parfenova, Angelina, et al.
Publicado: (2025)
por: Parfenova, Angelina, et al.
Publicado: (2025)
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
por: Refai, Dania, et al.
Publicado: (2025)
por: Refai, Dania, et al.
Publicado: (2025)
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: the Case of FrameNet Annotation
por: Belcavello, Frederico, et al.
Publicado: (2025)
por: Belcavello, Frederico, et al.
Publicado: (2025)
System Message Generation for User Preferences using Open-Source Models
por: Jeong, Minbyul, et al.
Publicado: (2025)
por: Jeong, Minbyul, et al.
Publicado: (2025)
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
por: Liu, Maoqi, et al.
Publicado: (2025)
por: Liu, Maoqi, et al.
Publicado: (2025)
Subjective Behaviors and Preferences in LLM: Language of Browsing
por: Sundaresan, Sai, et al.
Publicado: (2025)
por: Sundaresan, Sai, et al.
Publicado: (2025)
X-AMR Annotation Tool
por: Ahmed, Shafiuddin Rehan, et al.
Publicado: (2024)
por: Ahmed, Shafiuddin Rehan, et al.
Publicado: (2024)
EHR Interaction Between Patients and AI: NoteAid EHR Interaction
por: Zhang, Xiaocheng, et al.
Publicado: (2023)
por: Zhang, Xiaocheng, et al.
Publicado: (2023)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
por: Yeh, Samuel, et al.
Publicado: (2025)
por: Yeh, Samuel, et al.
Publicado: (2025)
Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings
por: Liu, Xuanqing, et al.
Publicado: (2025)
por: Liu, Xuanqing, et al.
Publicado: (2025)
RouteLLM: Learning to Route LLMs with Preference Data
por: Ong, Isaac, et al.
Publicado: (2024)
por: Ong, Isaac, et al.
Publicado: (2024)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
por: Parkar, Ritik Sachin, et al.
Publicado: (2024)
por: Parkar, Ritik Sachin, et al.
Publicado: (2024)
Personalized LLM Decoding via Contrasting Personal Preference
por: Bu, Hyungjune, et al.
Publicado: (2025)
por: Bu, Hyungjune, et al.
Publicado: (2025)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
por: Chen, Sirui, et al.
Publicado: (2025)
por: Chen, Sirui, et al.
Publicado: (2025)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
por: Zhang, Xuemiao, et al.
Publicado: (2025)
por: Zhang, Xuemiao, et al.
Publicado: (2025)
Aligning LLMs through Multi-perspective User Preference Ranking-based Feedback for Programming Question Answering
por: Yang, Hongyu, et al.
Publicado: (2024)
por: Yang, Hongyu, et al.
Publicado: (2024)
Learning to Summarize from LLM-generated Feedback
por: Song, Hwanjun, et al.
Publicado: (2024)
por: Song, Hwanjun, et al.
Publicado: (2024)
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
por: Wang, Zhilin, et al.
Publicado: (2025)
por: Wang, Zhilin, et al.
Publicado: (2025)
User-LLM: Efficient LLM Contextualization with User Embeddings
por: Ning, Lin, et al.
Publicado: (2024)
por: Ning, Lin, et al.
Publicado: (2024)
CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences
por: Pokharel, Rhitabrat, et al.
Publicado: (2025)
por: Pokharel, Rhitabrat, et al.
Publicado: (2025)
LRHP: Learning Representations for Human Preferences via Preference Pairs
por: Wang, Chenglong, et al.
Publicado: (2024)
por: Wang, Chenglong, et al.
Publicado: (2024)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
por: Liu, Wenxuan, et al.
Publicado: (2025)
por: Liu, Wenxuan, et al.
Publicado: (2025)
Ejemplares similares
-
Goal Alignment in LLM-Based User Simulators for Conversational AI
por: Mehri, Shuhaib, et al.
Publicado: (2025) -
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024) -
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
por: Wang, Yifan, et al.
Publicado: (2025) -
Arithmetic Reasoning with LLM: Prolog Generation & Permutation
por: Yang, Xiaocheng, et al.
Publicado: (2024) -
Aligning LLM Agents by Learning Latent Preference from User Edits
por: Gao, Ge, et al.
Publicado: (2024)