YIELD: A Large-Scale Dataset and Evaluation Framework for Information Elicitation Agents
Fuente:
arXiv
Saved in:
| Main Authors: | De Lima, Victor, Yang, Grace Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FrameRef: A Framing Dataset and Simulation Testbed for Modeling Bounded Rational Information Health
by: De Lima, Victor, et al.
Published: (2026)
by: De Lima, Victor, et al.
Published: (2026)
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference
by: Proebsting, Grace, et al.
Published: (2024)
by: Proebsting, Grace, et al.
Published: (2024)
Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference
by: Proebsting, Grace, et al.
Published: (2025)
by: Proebsting, Grace, et al.
Published: (2025)
SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
by: Jiang, Changhao, et al.
Published: (2025)
by: Jiang, Changhao, et al.
Published: (2025)
Eliciting Informative Text Evaluations with Large Language Models
by: Lu, Yuxuan, et al.
Published: (2024)
by: Lu, Yuxuan, et al.
Published: (2024)
JobHop: A Large-Scale Dataset of Career Trajectories
by: Johary, Iman, et al.
Published: (2025)
by: Johary, Iman, et al.
Published: (2025)
PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures
by: Lu, Yunan, et al.
Published: (2026)
by: Lu, Yunan, et al.
Published: (2026)
MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
by: Park, Ji-jun, et al.
Published: (2026)
by: Park, Ji-jun, et al.
Published: (2026)
Chain-of-Dictionary Prompting Elicits Translation in Large Language Models
by: Lu, Hongyuan, et al.
Published: (2023)
by: Lu, Hongyuan, et al.
Published: (2023)
Value of Information: A Framework for Human-Agent Communication
by: Dong, Yijiang River, et al.
Published: (2026)
by: Dong, Yijiang River, et al.
Published: (2026)
DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents
by: Chuang, Yun-Shiuan, et al.
Published: (2025)
by: Chuang, Yun-Shiuan, et al.
Published: (2025)
Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
by: Shen, Lei, et al.
Published: (2025)
by: Shen, Lei, et al.
Published: (2025)
MessIRve: A Large-Scale Spanish Information Retrieval Dataset
by: Valentini, Francisco, et al.
Published: (2024)
by: Valentini, Francisco, et al.
Published: (2024)
Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents
by: Lin, Xubo, et al.
Published: (2026)
by: Lin, Xubo, et al.
Published: (2026)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
by: Zhang, Zhaohan, et al.
Published: (2025)
by: Zhang, Zhaohan, et al.
Published: (2025)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
by: Nguyen, Long S. T., et al.
Published: (2025)
by: Nguyen, Long S. T., et al.
Published: (2025)
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
by: Liu, Maoqi, et al.
Published: (2025)
by: Liu, Maoqi, et al.
Published: (2025)
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
by: Zhang, Mozhi, et al.
Published: (2024)
by: Zhang, Mozhi, et al.
Published: (2024)
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
by: Liu, Yifei, et al.
Published: (2025)
by: Liu, Yifei, et al.
Published: (2025)
AutoElicit: Using Large Language Models for Expert Prior Elicitation in Predictive Modelling
by: Capstick, Alexander, et al.
Published: (2024)
by: Capstick, Alexander, et al.
Published: (2024)
Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks
by: Hegazy, Mahmood
Published: (2024)
by: Hegazy, Mahmood
Published: (2024)
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
by: Števaňák, David, et al.
Published: (2026)
by: Števaňák, David, et al.
Published: (2026)
Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit Reasoning
by: Li, Xintong, et al.
Published: (2025)
by: Li, Xintong, et al.
Published: (2025)
Executable Code Actions Elicit Better LLM Agents
by: Wang, Xingyao, et al.
Published: (2024)
by: Wang, Xingyao, et al.
Published: (2024)
Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models
by: Hu, Hanxu, et al.
Published: (2023)
by: Hu, Hanxu, et al.
Published: (2023)
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
KARRIEREWEGE: A Large Scale Career Path Prediction Dataset
by: Senger, Elena, et al.
Published: (2024)
by: Senger, Elena, et al.
Published: (2024)
Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
by: Qiu, Haoyi, et al.
Published: (2025)
by: Qiu, Haoyi, et al.
Published: (2025)
Eliciting Personality Traits in Large Language Models
by: Hilliard, Airlie, et al.
Published: (2024)
by: Hilliard, Airlie, et al.
Published: (2024)
FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models
by: Lin, Hongzhan, et al.
Published: (2025)
by: Lin, Hongzhan, et al.
Published: (2025)
Tip of the Tongue Query Elicitation for Simulated Evaluation
by: He, Yifan, et al.
Published: (2025)
by: He, Yifan, et al.
Published: (2025)
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
by: Stuart, Harry, et al.
Published: (2026)
by: Stuart, Harry, et al.
Published: (2026)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
by: Long, Zhuohan, et al.
Published: (2026)
by: Long, Zhuohan, et al.
Published: (2026)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
by: Xiong, Miao, et al.
Published: (2023)
by: Xiong, Miao, et al.
Published: (2023)
Eliciting Language Model Behaviors with Investigator Agents
by: Li, Xiang Lisa, et al.
Published: (2025)
by: Li, Xiang Lisa, et al.
Published: (2025)
Meta-Task Prompting Elicits Embeddings from Large Language Models
by: Lei, Yibin, et al.
Published: (2024)
by: Lei, Yibin, et al.
Published: (2024)
Eliciting the Priors of Large Language Models using Iterated In-Context Learning
by: Zhu, Jian-Qiao, et al.
Published: (2024)
by: Zhu, Jian-Qiao, et al.
Published: (2024)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Similar Items
-
FrameRef: A Framing Dataset and Simulation Testbed for Modeling Bounded Rational Information Health
by: De Lima, Victor, et al.
Published: (2026) -
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference
by: Proebsting, Grace, et al.
Published: (2024) -
Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference
by: Proebsting, Grace, et al.
Published: (2025) -
SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
by: Jiang, Changhao, et al.
Published: (2025) -
Eliciting Informative Text Evaluations with Large Language Models
by: Lu, Yuxuan, et al.
Published: (2024)