Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zongxia, Calvo-Bartolomé, Lorena, Hoyle, Alexander, Xu, Paiheng, Dima, Alden, Fung, Juan Francisco, Boyd-Graber, Jordan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
by: Hoyle, Alexander, et al.
Published: (2025)
by: Hoyle, Alexander, et al.
Published: (2025)
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
by: Gu, Feng, et al.
Published: (2025)
by: Gu, Feng, et al.
Published: (2025)
Labeled Interactive Topic Models
by: Seelman, Kyle, et al.
Published: (2023)
by: Seelman, Kyle, et al.
Published: (2023)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
Readme_AI: Dynamic Context Construction for Large Language Models
by: Vyas, Millie, et al.
Published: (2025)
by: Vyas, Millie, et al.
Published: (2025)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
by: Gor, Maharshi, et al.
Published: (2024)
by: Gor, Maharshi, et al.
Published: (2024)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
by: Mondal, Ishani, et al.
Published: (2024)
by: Mondal, Ishani, et al.
Published: (2024)
TopicGPT: A Prompt-based Topic Modeling Framework
by: Pham, Chau Minh, et al.
Published: (2023)
by: Pham, Chau Minh, et al.
Published: (2023)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
You've Changed: Detecting Modification of Black-Box Large Language Models
by: Dima, Alden, et al.
Published: (2025)
by: Dima, Alden, et al.
Published: (2025)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
by: Mondal, Ishani, et al.
Published: (2025)
by: Mondal, Ishani, et al.
Published: (2025)
Personalized Help for Optimizing Low-Skilled Users' Strategy
by: Gu, Feng, et al.
Published: (2024)
by: Gu, Feng, et al.
Published: (2024)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
Towards Understanding In-Context Learning with Contrastive Demonstrations and Saliency Maps
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA
by: Kabir, Tasnim, et al.
Published: (2026)
by: Kabir, Tasnim, et al.
Published: (2026)
Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
by: Cho, Eunjung, et al.
Published: (2025)
by: Cho, Eunjung, et al.
Published: (2025)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Self-Rewarding Vision-Language Model via Reasoning Decomposition
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
Multi-Hop Question Answering: When Can Humans Help, and Where do They Struggle?
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
by: Shu, Matthew, et al.
Published: (2024)
by: Shu, Matthew, et al.
Published: (2024)
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
by: Bilalpur, Maneesh, et al.
Published: (2025)
by: Bilalpur, Maneesh, et al.
Published: (2025)
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024)
by: Nelson, Elliot, et al.
Published: (2024)
Why Do Vision Language Models Struggle To Recognize Human Emotions?
by: Agarwal, Madhav, et al.
Published: (2026)
by: Agarwal, Madhav, et al.
Published: (2026)
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
by: Mao, Jiayuan, et al.
Published: (2023)
by: Mao, Jiayuan, et al.
Published: (2023)
Haystack observations of cs and ch3oh toward star forming regions
by: Evan Jordan
Published: (2008)
by: Evan Jordan
Published: (2008)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
by: Chochlakis, Georgios, et al.
Published: (2025)
by: Chochlakis, Georgios, et al.
Published: (2025)
Similar Items
-
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
by: Hoyle, Alexander, et al.
Published: (2025) -
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis
by: Li, Zongxia, et al.
Published: (2024) -
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
by: Gu, Feng, et al.
Published: (2025) -
Labeled Interactive Topic Models
by: Seelman, Kyle, et al.
Published: (2023) -
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)