Assessing and Verifying Task Utility in LLM-Powered Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Arabzadeh, Negar, Huo, Siqing, Mehta, Nikhil, Wu, Qinqyun, Wang, Chi, Awadallah, Ahmed, Clarke, Charles L. A., Kiseleva, Julia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
by: Mehta, Nikhil, et al.
Published: (2023)
by: Mehta, Nikhil, et al.
Published: (2023)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
by: Mohanty, Shrestha, et al.
Published: (2024)
by: Mohanty, Shrestha, et al.
Published: (2024)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task Adaptation
by: Dun, Chen, et al.
Published: (2023)
by: Dun, Chen, et al.
Published: (2023)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
by: Rosset, Corby, et al.
Published: (2024)
by: Rosset, Corby, et al.
Published: (2024)
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Reproducibility Study of LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Benchmarking Prompt Sensitivity in Large Language Models
by: Razavi, Amirhossein, et al.
Published: (2025)
by: Razavi, Amirhossein, et al.
Published: (2025)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
by: Wu, Yiran, et al.
Published: (2024)
by: Wu, Yiran, et al.
Published: (2024)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
by: Kanithi, Praveenkumar, et al.
Published: (2024)
by: Kanithi, Praveenkumar, et al.
Published: (2024)
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions
by: Hu, Yiran, et al.
Published: (2026)
by: Hu, Yiran, et al.
Published: (2026)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
by: Li, Xiaonan, et al.
Published: (2023)
by: Li, Xiaonan, et al.
Published: (2023)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
by: Su, Zhe, et al.
Published: (2024)
by: Su, Zhe, et al.
Published: (2024)
VerAs: Verify then Assess STEM Lab Reports
by: Atil, Berk, et al.
Published: (2024)
by: Atil, Berk, et al.
Published: (2024)
OmniParser for Pure Vision Based GUI Agent
by: Lu, Yadong, et al.
Published: (2024)
by: Lu, Yadong, et al.
Published: (2024)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
by: Sharma, Nikhil, et al.
Published: (2024)
by: Sharma, Nikhil, et al.
Published: (2024)
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
by: Zhang, Ran, et al.
Published: (2026)
by: Zhang, Ran, et al.
Published: (2026)
Optimizing Pretraining Data Mixtures with LLM-Estimated Utility
by: Held, William, et al.
Published: (2025)
by: Held, William, et al.
Published: (2025)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
by: Rosset, Corby, et al.
Published: (2024)
by: Rosset, Corby, et al.
Published: (2024)
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
Automatic Pair Construction for Contrastive Post-training
by: Xu, Canwen, et al.
Published: (2023)
by: Xu, Canwen, et al.
Published: (2023)
Performance of a large language model-Artificial Intelligence based chatbot for counseling patients with sexually transmitted infections and genital diseases
by: Mehta, Nikhil, et al.
Published: (2024)
by: Mehta, Nikhil, et al.
Published: (2024)
Systematic Analysis of LLM Contributions to Planning: Solver, Verifier, Heuristic
by: Li, Haoming, et al.
Published: (2024)
by: Li, Haoming, et al.
Published: (2024)
Similar Items
-
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024) -
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
by: Mehta, Nikhil, et al.
Published: (2023) -
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025) -
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025) -
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
by: Mohanty, Shrestha, et al.
Published: (2024)