IDRBench: Interactive Deep Research Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yingchaojie, Huang, Qiang, Xie, Xiaoya, Yang, Zhaorui, Yu, Jun, Chen, Wei, Tung, Anthony K. H. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
by: Tian, Haoyu, et al.
Published: (2026)
by: Tian, Haoyu, et al.
Published: (2026)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
InsightLens: Augmenting LLM-Powered Data Analysis with Interactive Insight Management and Navigation
by: Weng, Luoxuan, et al.
Published: (2024)
by: Weng, Luoxuan, et al.
Published: (2024)
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
by: Qiu, Rui, et al.
Published: (2025)
by: Qiu, Rui, et al.
Published: (2025)
SQLucid: Grounding Natural Language Database Queries with Interactive Explanations
by: Tian, Yuan, et al.
Published: (2024)
by: Tian, Yuan, et al.
Published: (2024)
ARAIDA: Analogical Reasoning-Augmented Interactive Data Annotation
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
Generating Educational Materials with Different Levels of Readability using LLMs
by: Huang, Chieh-Yang, et al.
Published: (2024)
by: Huang, Chieh-Yang, et al.
Published: (2024)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
by: Nguyen, Hy, et al.
Published: (2024)
by: Nguyen, Hy, et al.
Published: (2024)
"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews
by: Wan, Ruyuan, et al.
Published: (2026)
by: Wan, Ruyuan, et al.
Published: (2026)
CultiVerse: Towards Cross-Cultural Understanding for Paintings with Large Language Model
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Learning Next Action Predictors from Human-Computer Interaction
by: Shaikh, Omar, et al.
Published: (2026)
by: Shaikh, Omar, et al.
Published: (2026)
Deep Learning and Machine Learning -- Natural Language Processing: From Theory to Application
by: Chen, Keyu, et al.
Published: (2024)
by: Chen, Keyu, et al.
Published: (2024)
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration
by: Pan, Bo, et al.
Published: (2024)
by: Pan, Bo, et al.
Published: (2024)
Effects of Collaboration on the Performance of Interactive Theme Discovery Systems
by: Chen, Alvin Po-Chun, et al.
Published: (2024)
by: Chen, Alvin Po-Chun, et al.
Published: (2024)
NARRA-Gym for Evaluating Interactive Narrative Agents
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
AgentA/B: Automated and Scalable Web A/BTesting with Interactive LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
MONAL: Model Autophagy Analysis for Modeling Human-AI Interactions
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
A Design Space for Intelligent and Interactive Writing Assistants
by: Lee, Mina, et al.
Published: (2024)
by: Lee, Mina, et al.
Published: (2024)
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset
by: Wang, Fen, et al.
Published: (2025)
by: Wang, Fen, et al.
Published: (2025)
Exploring the Efficacy of Large Language Models in Summarizing Mental Health Counseling Sessions: A Benchmark Study
by: Adhikary, Prottay Kumar, et al.
Published: (2024)
by: Adhikary, Prottay Kumar, et al.
Published: (2024)
Automatic Macro Mining from Interaction Traces at Scale
by: Huang, Forrest, et al.
Published: (2023)
by: Huang, Forrest, et al.
Published: (2023)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research
by: Pi, Xinyu, et al.
Published: (2025)
by: Pi, Xinyu, et al.
Published: (2025)
Dialogue Act Patterns in GenAI-Mediated L2 Oral Practice: A Sequential Analysis of Learner-Chatbot Interactions
by: He, Liqun, et al.
Published: (2026)
by: He, Liqun, et al.
Published: (2026)
Measuring Mental Health Variables in Computational Research: Toward Validated, Dimensional, and Transdiagnostic Approaches
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs
by: Xiang, Yurui, et al.
Published: (2026)
by: Xiang, Yurui, et al.
Published: (2026)
Interactive Recommendation Agent with Active User Commands
by: Tang, Jiakai, et al.
Published: (2025)
by: Tang, Jiakai, et al.
Published: (2025)
Deceptive Patterns of Intelligent and Interactive Writing Assistants
by: Benharrak, Karim, et al.
Published: (2024)
by: Benharrak, Karim, et al.
Published: (2024)
LIVE: LaTex Interactive Visual Editing
by: Lin, Jinwei
Published: (2024)
by: Lin, Jinwei
Published: (2024)
Modeling Distinct Human Interaction in Web Agents
by: Huq, Faria, et al.
Published: (2026)
by: Huq, Faria, et al.
Published: (2026)
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
by: Yang, Jackie Junrui, et al.
Published: (2023)
by: Yang, Jackie Junrui, et al.
Published: (2023)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
by: Yu, Yeyong, et al.
Published: (2024)
by: Yu, Yeyong, et al.
Published: (2024)
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
by: Kim, Seon Gyeom, et al.
Published: (2025)
by: Kim, Seon Gyeom, et al.
Published: (2025)
Spoken Language Interaction with Robots: Research Issues and Recommendations, Report from the NSF Future Directions Workshop
by: Marge, Matthew, et al.
Published: (2020)
by: Marge, Matthew, et al.
Published: (2020)
Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows
by: Balashov, Yuri, et al.
Published: (2026)
by: Balashov, Yuri, et al.
Published: (2026)
Similar Items
-
RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
by: Tian, Haoyu, et al.
Published: (2026) -
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024) -
InsightLens: Augmenting LLM-Powered Data Analysis with Interactive Insight Management and Navigation
by: Weng, Luoxuan, et al.
Published: (2024) -
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
by: Tang, Yinghao, et al.
Published: (2026) -
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)