IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Mohanty, Shrestha, Arabzadeh, Negar, Tupini, Andrea, Sun, Yuxuan, Skrynnik, Alexey, Zholus, Artem, Côté, Marc-Alexandre, Kiseleva, Julia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAMAR: Continuous Actions Multi-Agent Routing
by: Pshenitsyn, Artem, et al.
Published: (2025)
by: Pshenitsyn, Artem, et al.
Published: (2025)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Assessing and Verifying Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Continuous Histogram Loss: Beyond Neural Similarity
by: Zholus, Artem, et al.
Published: (2020)
by: Zholus, Artem, et al.
Published: (2020)
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
by: Cherepanov, Egor, et al.
Published: (2024)
by: Cherepanov, Egor, et al.
Published: (2024)
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
CoRL-MPPI: Enhancing MPPI With Learnable Behaviours For Efficient And Provably-Safe Multi-Robot Collision Avoidance
by: Dergachev, Stepan, et al.
Published: (2025)
by: Dergachev, Stepan, et al.
Published: (2025)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at Scale
by: Andreychuk, Anton, et al.
Published: (2024)
by: Andreychuk, Anton, et al.
Published: (2024)
Advancing Learnable Multi-Agent Pathfinding Solvers with Active Fine-Tuning
by: Andreychuk, Anton, et al.
Published: (2025)
by: Andreychuk, Anton, et al.
Published: (2025)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Optimal Dataset Size for Recommender Systems: Evaluating Algorithms' Performance via Downsampling
by: Arabzadeh, Ardalan
Published: (2025)
by: Arabzadeh, Ardalan
Published: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning
by: Volovikova, Zoya, et al.
Published: (2026)
by: Volovikova, Zoya, et al.
Published: (2026)
DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Gravitational Redshift and Variable Speed of Light: An Alternative to Spacetime Curvature
by: Skrynnik, Sergey
Published: (2025)
by: Skrynnik, Sergey
Published: (2025)
AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback
by: Tupini, Andrea, et al.
Published: (2026)
by: Tupini, Andrea, et al.
Published: (2026)
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Natural Language Query to Configuration for Retrieval Agents
by: Pan, Melissa Z., et al.
Published: (2026)
by: Pan, Melissa Z., et al.
Published: (2026)
POGEMA: A Benchmark Platform for Cooperative Multi-Agent Pathfinding
by: Skrynnik, Alexey, et al.
Published: (2024)
by: Skrynnik, Alexey, et al.
Published: (2024)
Revisiting Tree Search for LLMs: Gumbel and Sequential Halving for Budget-Scalable Reasoning
by: Ugadiarov, Leonid, et al.
Published: (2026)
by: Ugadiarov, Leonid, et al.
Published: (2026)
AgentStudio: A Toolkit for Building General Virtual Agents
by: Zheng, Longtao, et al.
Published: (2024)
by: Zheng, Longtao, et al.
Published: (2024)
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
by: Mehta, Nikhil, et al.
Published: (2023)
by: Mehta, Nikhil, et al.
Published: (2023)
Instruction Following with Goal-Conditioned Reinforcement Learning in Virtual Environments
by: Volovikova, Zoya, et al.
Published: (2024)
by: Volovikova, Zoya, et al.
Published: (2024)
MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning
by: Nesterova, Maria, et al.
Published: (2026)
by: Nesterova, Maria, et al.
Published: (2026)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
by: Ebrahimi, Sajad, et al.
Published: (2025)
by: Ebrahimi, Sajad, et al.
Published: (2025)
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
by: Bigdeli, Amin, et al.
Published: (2024)
by: Bigdeli, Amin, et al.
Published: (2024)
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024)
by: Alaofi, Marwah, et al.
Published: (2024)
The study of histopathological damage in heart and bulbus arteriosus in rainbow trout
by: Arabzadeh, Payam
Published: (2009)
by: Arabzadeh, Payam
Published: (2009)
Green Recommender Systems: Optimizing Dataset Size for Energy-Efficient Algorithm Performance
by: Arabzadeh, Ardalan, et al.
Published: (2024)
by: Arabzadeh, Ardalan, et al.
Published: (2024)
SplitLight: An Exploratory Toolkit for Recommender Systems Datasets and Splits
by: Volodkevich, Anna, et al.
Published: (2026)
by: Volodkevich, Anna, et al.
Published: (2026)
Sub-goal Distillation: A Method to Improve Small Language Agents
by: Hashemzadeh, Maryam, et al.
Published: (2024)
by: Hashemzadeh, Maryam, et al.
Published: (2024)
Software Implementation of an Algorithm for Solving a Dynamic Problem of Optimal Set Partitioning Under Uncertainty
by: Kiseleva, E. M., et al.
Published: (2025)
by: Kiseleva, E. M., et al.
Published: (2025)
Similar Items
-
CAMAR: Continuous Actions Multi-Agent Routing
by: Pshenitsyn, Artem, et al.
Published: (2025) -
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024) -
Assessing and Verifying Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024) -
Continuous Histogram Loss: Beyond Neural Similarity
by: Zholus, Artem, et al.
Published: (2020) -
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
by: Cherepanov, Egor, et al.
Published: (2024)