Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Malaviya, Chaitanya, Chang, Joseph Chee, Roth, Dan, Iyyer, Mohit, Yatskar, Mark, Lo, Kyle |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
by: Malaviya, Chaitanya, et al.
Published: (2023)
by: Malaviya, Chaitanya, et al.
Published: (2023)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
by: Yifei, Li S., et al.
Published: (2025)
by: Yifei, Li S., et al.
Published: (2025)
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
by: Bharadwaj, Anirudh, et al.
Published: (2025)
by: Bharadwaj, Anirudh, et al.
Published: (2025)
ExpertQA: Expert-Curated Questions and Attributed Answers
by: Malaviya, Chaitanya, et al.
Published: (2023)
by: Malaviya, Chaitanya, et al.
Published: (2023)
A Simple Joint Model for Improved Contextual Neural Lemmatization
by: Malaviya, Chaitanya, et al.
Published: (2019)
by: Malaviya, Chaitanya, et al.
Published: (2019)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023)
by: Chang, Yapei, et al.
Published: (2023)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
On Reference (In-)Determinacy in Natural Language Inference
by: Chen, Sihao, et al.
Published: (2025)
by: Chen, Sihao, et al.
Published: (2025)
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024)
by: Karpinska, Marzena, et al.
Published: (2024)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
by: Samuel, Vinay, et al.
Published: (2026)
by: Samuel, Vinay, et al.
Published: (2026)
FABLES: Evaluating faithfulness and content selection in book-length summarization
by: Kim, Yekyung, et al.
Published: (2024)
by: Kim, Yekyung, et al.
Published: (2024)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
by: Jafari, Nazanin, et al.
Published: (2026)
by: Jafari, Nazanin, et al.
Published: (2026)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
DOLOMITES: Domain-Specific Long-Form Methodical Tasks
by: Malaviya, Chaitanya, et al.
Published: (2024)
by: Malaviya, Chaitanya, et al.
Published: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
by: Pham, Chau Minh, et al.
Published: (2024)
by: Pham, Chau Minh, et al.
Published: (2024)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
by: Qu, Renyi, et al.
Published: (2024)
by: Qu, Renyi, et al.
Published: (2024)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
by: Mukhopadhyay, Srija, et al.
Published: (2024)
by: Mukhopadhyay, Srija, et al.
Published: (2024)
Intent-Aware Schema Generation And Refinement For Literature Review Tables
by: Padmakumar, Vishakh, et al.
Published: (2025)
by: Padmakumar, Vishakh, et al.
Published: (2025)
ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
by: Newman, Benjamin, et al.
Published: (2024)
by: Newman, Benjamin, et al.
Published: (2024)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
by: Yoran, Ori, et al.
Published: (2024)
by: Yoran, Ori, et al.
Published: (2024)
Localizing and Mitigating Errors in Long-form Question Answering
by: Sachdeva, Rachneet, et al.
Published: (2024)
by: Sachdeva, Rachneet, et al.
Published: (2024)
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
EditLens: Quantifying the Extent of AI Editing in Text
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
by: Chang, Yapei, et al.
Published: (2026)
by: Chang, Yapei, et al.
Published: (2026)
TopicGPT: A Prompt-based Topic Modeling Framework
by: Pham, Chau Minh, et al.
Published: (2023)
by: Pham, Chau Minh, et al.
Published: (2023)
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Calibrating Large Language Models with Sample Consistency
by: Lyu, Qing, et al.
Published: (2024)
by: Lyu, Qing, et al.
Published: (2024)
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Whose story is it? Personalizing story generation by inferring author styles
by: Kumar, Nischal Ashok, et al.
Published: (2025)
by: Kumar, Nischal Ashok, et al.
Published: (2025)
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025)
by: Rajendhran, Rishanth, et al.
Published: (2025)
Bridging the Gap: Transforming Natural Language Questions into SQL Queries via Abstract Query Pattern and Contextual Schema Markup
by: Kong, Yonghui, et al.
Published: (2025)
by: Kong, Yonghui, et al.
Published: (2025)
Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance
by: Huang, Yunchong, et al.
Published: (2026)
by: Huang, Yunchong, et al.
Published: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Similar Items
-
What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
by: Malaviya, Chaitanya, et al.
Published: (2023) -
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
by: Yifei, Li S., et al.
Published: (2025) -
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
by: Bharadwaj, Anirudh, et al.
Published: (2025) -
ExpertQA: Expert-Curated Questions and Attributed Answers
by: Malaviya, Chaitanya, et al.
Published: (2023) -
A Simple Joint Model for Improved Contextual Neural Lemmatization
by: Malaviya, Chaitanya, et al.
Published: (2019)