A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Balepur, Nishant, Shu, Matthew, Sung, Yoo Yeon, Goldfarb-Tarrant, Seraphina, Feng, Shi, Yang, Fumeng, Rudinger, Rachel, Boyd-Graber, Jordan Lee |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
par: Shu, Matthew, et autres
Publié: (2024)
par: Shu, Matthew, et autres
Publié: (2024)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
par: Balepur, Nishant, et autres
Publié: (2023)
par: Balepur, Nishant, et autres
Publié: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
par: Chen, Hongyu, et autres
Publié: (2025)
par: Chen, Hongyu, et autres
Publié: (2025)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
par: Srikanth, Neha, et autres
Publié: (2026)
par: Srikanth, Neha, et autres
Publié: (2026)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
par: Sung, Yoo Yeon, et autres
Publié: (2024)
par: Sung, Yoo Yeon, et autres
Publié: (2024)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
par: Aakanksha, et autres
Publié: (2024)
par: Aakanksha, et autres
Publié: (2024)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
par: Seshadri, Preethi, et autres
Publié: (2026)
par: Seshadri, Preethi, et autres
Publié: (2026)
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
par: Balepur, Nishant, et autres
Publié: (2026)
par: Balepur, Nishant, et autres
Publié: (2026)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
par: Balepur, Nishant, et autres
Publié: (2026)
par: Balepur, Nishant, et autres
Publié: (2026)
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
par: Seshadri, Preethi, et autres
Publié: (2025)
par: Seshadri, Preethi, et autres
Publié: (2025)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
par: Balepur, Nishant, et autres
Publié: (2026)
par: Balepur, Nishant, et autres
Publié: (2026)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
par: Palta, Shramay, et autres
Publié: (2024)
par: Palta, Shramay, et autres
Publié: (2024)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
par: Gor, Maharshi, et autres
Publié: (2026)
par: Gor, Maharshi, et autres
Publié: (2026)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
par: Rajani, Neel, et autres
Publié: (2025)
par: Rajani, Neel, et autres
Publié: (2025)
MultiContrievers: Analysis of Dense Retrieval Representations
par: Goldfarb-Tarrant, Seraphina, et autres
Publié: (2024)
par: Goldfarb-Tarrant, Seraphina, et autres
Publié: (2024)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
par: Sung, Yoo Yeon, et autres
Publié: (2024)
par: Sung, Yoo Yeon, et autres
Publié: (2024)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
par: Sung, Yoo Yeon, et autres
Publié: (2025)
par: Sung, Yoo Yeon, et autres
Publié: (2025)
Personalized Help for Optimizing Low-Skilled Users' Strategy
par: Gu, Feng, et autres
Publié: (2024)
par: Gu, Feng, et autres
Publié: (2024)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
par: Aakanksha, et autres
Publié: (2024)
par: Aakanksha, et autres
Publié: (2024)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
par: Srikanth, Neha, et autres
Publié: (2023)
par: Srikanth, Neha, et autres
Publié: (2023)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
par: Srikanth, Neha, et autres
Publié: (2025)
par: Srikanth, Neha, et autres
Publié: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
par: Orgad, Hadas, et autres
Publié: (2026)
par: Orgad, Hadas, et autres
Publié: (2026)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
par: Ki, Dayeon, et autres
Publié: (2026)
par: Ki, Dayeon, et autres
Publié: (2026)
Labeled Interactive Topic Models
par: Seelman, Kyle, et autres
Publié: (2023)
par: Seelman, Kyle, et autres
Publié: (2023)
Read Any Good Books Lately? Helping Patrons Find What They Want.
par: Chelton, Mary K.
Publié: (1993)
par: Chelton, Mary K.
Publié: (1993)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
par: Gu, Feng, et autres
Publié: (2025)
par: Gu, Feng, et autres
Publié: (2025)
Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs
par: Sarkar, Rupak, et autres
Publié: (2025)
par: Sarkar, Rupak, et autres
Publié: (2025)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
par: Si, Chenglei, et autres
Publié: (2023)
par: Si, Chenglei, et autres
Publié: (2023)
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
par: Palta, Shramay, et autres
Publié: (2025)
par: Palta, Shramay, et autres
Publié: (2025)
Documents similaires
-
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
par: Balepur, Nishant, et autres
Publié: (2025) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
par: Balepur, Nishant, et autres
Publié: (2025) -
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
par: Balepur, Nishant, et autres
Publié: (2024) -
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
par: Balepur, Nishant, et autres
Publié: (2024) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
par: Balepur, Nishant, et autres
Publié: (2024)