Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
Fuente:
arXiv
Saved in:
| Main Authors: | Balepur, Nishant, Hamada, Malachi, Kishore, Varsha, Feldman, Sergey, Singh, Amanpreet, Siangliulue, Pao, Chang, Joseph Chee, Choi, Eunsol, Boyd-Graber, Jordan Lee, Naik, Aakanksha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
by: Balepur, Nishant, et al.
Published: (2026)
by: Balepur, Nishant, et al.
Published: (2026)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
by: Shu, Matthew, et al.
Published: (2024)
by: Shu, Matthew, et al.
Published: (2024)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
by: Hwang, Jena D., et al.
Published: (2026)
by: Hwang, Jena D., et al.
Published: (2026)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
LitPivot: Developing Well-Situated Research Ideas Through Dynamic Contextualization and Critique within the Literature Landscape
by: Kambhamettu, Hita, et al.
Published: (2026)
by: Kambhamettu, Hita, et al.
Published: (2026)
If You Want the University to Change, Don't Theorise—Organise!
by: Sol Gamsu
Published: (2025)
by: Sol Gamsu
Published: (2025)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024)
by: Johnson, Daniel D., et al.
Published: (2024)
“They Don't Want to Know; They Don't Want to Hear”: Social Distance Between Leaders and Low‐Income Community Members in a Rural Indiana Community☆
by: Steven Tuttle, et al.
Published: (2025)
by: Steven Tuttle, et al.
Published: (2025)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
by: Newman, Benjamin, et al.
Published: (2024)
by: Newman, Benjamin, et al.
Published: (2024)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
by: Balepur, Nishant, et al.
Published: (2026)
by: Balepur, Nishant, et al.
Published: (2026)
When You Don't Know the Answer, Say So
Published: (2024)
Published: (2024)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Don't Use the Names Students Want
by: Rebecca Weaver
Published: (2025)
by: Rebecca Weaver
Published: (2025)
Improving Attributed Long-form Question Answering with Intent Awareness
by: Zhao, Xinran, et al.
Published: (2026)
by: Zhao, Xinran, et al.
Published: (2026)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
by: Pal, Arka, et al.
Published: (2025)
by: Pal, Arka, et al.
Published: (2025)
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
by: Khan, Imran
Published: (2025)
by: Khan, Imran
Published: (2025)
You Don't Need Public Tests to Generate Correct Code
by: Silva, Kaushitha, et al.
Published: (2026)
by: Silva, Kaushitha, et al.
Published: (2026)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
How to Computerize Your Serials and Periodicals When You Don't Know How
by: Matthews, Mary, et al.
Published: (1970)
by: Matthews, Mary, et al.
Published: (1970)
Omakase: proactive assistance with actionable suggestions for evolving scientific research projects
by: Siangliulue, Pao, et al.
Published: (2026)
by: Siangliulue, Pao, et al.
Published: (2026)
Papers-to-Posts: Supporting Detailed Long-Document Summarization with an Interactive LLM-Powered Source Outline
by: Radensky, Marissa, et al.
Published: (2024)
by: Radensky, Marissa, et al.
Published: (2024)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Imagining What We Don't Know
by: Samuels, Lisa
Published: (2026)
by: Samuels, Lisa
Published: (2026)
Imagining What We Don't Know
by: Samuels, Lisa
Published: (2026)
by: Samuels, Lisa
Published: (2026)
You See It, They Don't: An Exploratory Study of User-to-User Variation in Instagram Comments
by: Nutakki, Brahmani, et al.
Published: (2026)
by: Nutakki, Brahmani, et al.
Published: (2026)
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
by: Bian, Yutong, et al.
Published: (2025)
by: Bian, Yutong, et al.
Published: (2025)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
by: Goloburda, Maiya, et al.
Published: (2026)
by: Goloburda, Maiya, et al.
Published: (2026)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
Similar Items
-
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
by: Balepur, Nishant, et al.
Published: (2026) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025) -
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
by: Shu, Matthew, et al.
Published: (2024) -
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
by: Hwang, Jena D., et al.
Published: (2026) -
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
by: Balepur, Nishant, et al.
Published: (2025)