Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Arabzadeh, Negar, Kiseleva, Julia, Wu, Qingyun, Wang, Chi, Awadallah, Ahmed, Dibia, Victor, Fourney, Adam, Clarke, Charles |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing and Verifying Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Challenges in Human-Agent Communication
by: Bansal, Gagan, et al.
Published: (2024)
by: Bansal, Gagan, et al.
Published: (2024)
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
by: Mohanty, Shrestha, et al.
Published: (2024)
by: Mohanty, Shrestha, et al.
Published: (2024)
Interactive Debugging and Steering of Multi-Agent AI Systems
by: Epperson, Will, et al.
Published: (2025)
by: Epperson, Will, et al.
Published: (2025)
AutoGen Studio: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems
by: Dibia, Victor, et al.
Published: (2024)
by: Dibia, Victor, et al.
Published: (2024)
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
by: Mehta, Nikhil, et al.
Published: (2023)
by: Mehta, Nikhil, et al.
Published: (2023)
Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
by: Fourney, Adam, et al.
Published: (2024)
by: Fourney, Adam, et al.
Published: (2024)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
by: Bigdeli, Amin, et al.
Published: (2024)
by: Bigdeli, Amin, et al.
Published: (2024)
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024)
by: Alaofi, Marwah, et al.
Published: (2024)
Optimizing Sequential Multi-Step Tasks with Parallel LLM Agents
by: Zhang, Enhao, et al.
Published: (2025)
by: Zhang, Enhao, et al.
Published: (2025)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
A Reproducibility Study of LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2023)
by: Mozannar, Hussein, et al.
Published: (2023)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Natural Language Query to Configuration for Retrieval Agents
by: Pan, Melissa Z., et al.
Published: (2026)
by: Pan, Melissa Z., et al.
Published: (2026)
Magentic-UI: Towards Human-in-the-loop Agentic Systems
by: Mozannar, Hussein, et al.
Published: (2025)
by: Mozannar, Hussein, et al.
Published: (2025)
Accurate Measures of Vaccination and Concerns of Vaccine Holdouts from Web Search Logs
by: Chang, Serina, et al.
Published: (2023)
by: Chang, Serina, et al.
Published: (2023)
Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task Adaptation
by: Dun, Chen, et al.
Published: (2023)
by: Dun, Chen, et al.
Published: (2023)
Beyond Utility: Evaluating LLM as Recommender
by: Jiang, Chumeng, et al.
Published: (2024)
by: Jiang, Chumeng, et al.
Published: (2024)
OmniParser for Pure Vision Based GUI Agent
by: Lu, Yadong, et al.
Published: (2024)
by: Lu, Yadong, et al.
Published: (2024)
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
by: Wu, Yiran, et al.
Published: (2024)
by: Wu, Yiran, et al.
Published: (2024)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
by: Ebrahimi, Sajad, et al.
Published: (2025)
by: Ebrahimi, Sajad, et al.
Published: (2025)
Optimal Dataset Size for Recommender Systems: Evaluating Algorithms' Performance via Downsampling
by: Arabzadeh, Ardalan
Published: (2025)
by: Arabzadeh, Ardalan
Published: (2025)
The study of histopathological damage in heart and bulbus arteriosus in rainbow trout
by: Arabzadeh, Payam
Published: (2009)
by: Arabzadeh, Payam
Published: (2009)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
by: Zhang, Shaokun, et al.
Published: (2025)
by: Zhang, Shaokun, et al.
Published: (2025)
Towards LLM-Empowered Knowledge Tracing via LLM-Student Hierarchical Behavior Alignment in Hyperbolic Space
by: Fu, Xingcheng, et al.
Published: (2026)
by: Fu, Xingcheng, et al.
Published: (2026)
The Aleph & Other Metaphors for Image Generation
by: Ramos, Gonzalo, et al.
Published: (2024)
by: Ramos, Gonzalo, et al.
Published: (2024)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Similar Items
-
Assessing and Verifying Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024) -
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025) -
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025) -
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024) -
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024)