DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Balepur, Nishant, Hamada, Malachi, Kishore, Varsha, Feldman, Sergey, Singh, Amanpreet, Siangliulue, Pao, Chang, Joseph Chee, Rudinger, Rachel, Choi, Eunsol, Boyd-Graber, Jordan Lee, Downey, Doug, Naik, Aakanksha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
von: Balepur, Nishant, et al.
Veröffentlicht: (2023)
von: Balepur, Nishant, et al.
Veröffentlicht: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Omakase: proactive assistance with actionable suggestions for evolving scientific research projects
von: Siangliulue, Pao, et al.
Veröffentlicht: (2026)
von: Siangliulue, Pao, et al.
Veröffentlicht: (2026)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
von: Shu, Matthew, et al.
Veröffentlicht: (2024)
von: Shu, Matthew, et al.
Veröffentlicht: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2026)
von: Srikanth, Neha, et al.
Veröffentlicht: (2026)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
Intent-Aware Schema Generation And Refinement For Literature Review Tables
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2025)
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
LitPivot: Developing Well-Situated Research Ideas Through Dynamic Contextualization and Critique within the Literature Landscape
von: Kambhamettu, Hita, et al.
Veröffentlicht: (2026)
von: Kambhamettu, Hita, et al.
Veröffentlicht: (2026)
CARE: Extracting Experimental Findings From Clinical Literature
von: Naik, Aakanksha, et al.
Veröffentlicht: (2023)
von: Naik, Aakanksha, et al.
Veröffentlicht: (2023)
TOPICAL: TOPIC Pages AutomagicaLly
von: Giorgi, John, et al.
Veröffentlicht: (2024)
von: Giorgi, John, et al.
Veröffentlicht: (2024)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
von: Newman, Benjamin, et al.
Veröffentlicht: (2024)
von: Newman, Benjamin, et al.
Veröffentlicht: (2024)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Improving Attributed Long-form Question Answering with Intent Awareness
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)
Papers-to-Posts: Supporting Detailed Long-Document Summarization with an Interactive LLM-Powered Source Outline
von: Radensky, Marissa, et al.
Veröffentlicht: (2024)
von: Radensky, Marissa, et al.
Veröffentlicht: (2024)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2023)
von: Srikanth, Neha, et al.
Veröffentlicht: (2023)
Cocoa: Co-Planning and Co-Execution with AI Agents
von: Feng, K. J. Kevin, et al.
Veröffentlicht: (2024)
von: Feng, K. J. Kevin, et al.
Veröffentlicht: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Ai2 Scholar QA: Organized Literature Synthesis with Attribution
von: Singh, Amanpreet, et al.
Veröffentlicht: (2025)
von: Singh, Amanpreet, et al.
Veröffentlicht: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
Labeled Interactive Topic Models
von: Seelman, Kyle, et al.
Veröffentlicht: (2023)
von: Seelman, Kyle, et al.
Veröffentlicht: (2023)
LORD OF THE FLIES: POLLINATION OF DRACULA ORCHIDS
von: Lorena Endara
Veröffentlicht: (2010)
von: Lorena Endara
Veröffentlicht: (2010)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2024)
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2024)
Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review
von: Fok, Raymond, et al.
Veröffentlicht: (2025)
von: Fok, Raymond, et al.
Veröffentlicht: (2025)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2025)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
von: Hoyle, Alexander, et al.
Veröffentlicht: (2025)
von: Hoyle, Alexander, et al.
Veröffentlicht: (2025)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
von: Gor, Maharshi, et al.
Veröffentlicht: (2024)
von: Gor, Maharshi, et al.
Veröffentlicht: (2024)
On-the-fly Definition Augmentation of LLMs for Biomedical NER
von: Munnangi, Monica, et al.
Veröffentlicht: (2024)
von: Munnangi, Monica, et al.
Veröffentlicht: (2024)
Literature-Grounded Novelty Assessment of Scientific Ideas
von: Shahid, Simra, et al.
Veröffentlicht: (2025)
von: Shahid, Simra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2026) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
von: Balepur, Nishant, et al.
Veröffentlicht: (2025) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024) -
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024) -
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)