DRBench: A Realistic Benchmark for Enterprise Deep Research
Fuente:
arXiv
Saved in:
| Main Authors: | Abaskohi, Amirhossein, Chen, Tianyi, Muñoz-Mármol, Miguel, Fox, Curtis, Ramesh, Amrutha Varshini, Marcotte, Étienne, Lù, Xing Han, Chapados, Nicolas, Gella, Spandana, West, Peter, Carenini, Giuseppe, Pal, Christopher, Drouin, Alexandre, Laradji, Issam H. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
BlockLLM: Memory-Efficient Adaptation of LLMs by Selecting and Optimizing the Right Coordinate Blocks
by: Ramesh, Amrutha Varshini, et al.
Published: (2024)
by: Ramesh, Amrutha Varshini, et al.
Published: (2024)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
by: Ashok, Arjun, et al.
Published: (2023)
by: Ashok, Arjun, et al.
Published: (2023)
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction
by: Abaskohi, Amirhossein, et al.
Published: (2026)
by: Abaskohi, Amirhossein, et al.
Published: (2026)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
CEMTM: Contextual Embedding-based Multimodal Topic Modeling
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
Improving Neural Topic Modeling with Semantically-Grounded Soft Label Distributions
by: Li, Raymond, et al.
Published: (2026)
by: Li, Raymond, et al.
Published: (2026)
Dr-CiK: A Testbed for Foresight-Driven Agents
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
by: Ashok, Arjun, et al.
Published: (2025)
by: Ashok, Arjun, et al.
Published: (2025)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
by: Drouin, Alexandre, et al.
Published: (2024)
by: Drouin, Alexandre, et al.
Published: (2024)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
by: Salamatian, Ali, et al.
Published: (2025)
by: Salamatian, Ali, et al.
Published: (2025)
Context is Key: A Benchmark for Forecasting with Essential Textual Information
by: Williams, Andrew Robert, et al.
Published: (2024)
by: Williams, Andrew Robert, et al.
Published: (2024)
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization
by: Sahu, Gaurav, et al.
Published: (2023)
by: Sahu, Gaurav, et al.
Published: (2023)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
by: Sadeghi, Pouya, et al.
Published: (2024)
by: Sadeghi, Pouya, et al.
Published: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
SSLR: A Semi-Supervised Learning Method for Isolated Sign Language Recognition
by: Algafri, Hasan, et al.
Published: (2025)
by: Algafri, Hasan, et al.
Published: (2025)
IntentGPT: Few-shot Intent Discovery with Large Language Models
by: Rodriguez, Juan A., et al.
Published: (2024)
by: Rodriguez, Juan A., et al.
Published: (2024)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
Overcoming the Modality Gap in Context-Aided Forecasting
by: Zheng, Vincent Zhihao, et al.
Published: (2026)
by: Zheng, Vincent Zhihao, et al.
Published: (2026)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
Improving Language Models with Intentional Analysis
by: Yin, Yuwei, et al.
Published: (2025)
by: Yin, Yuwei, et al.
Published: (2025)
BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
by: Li, Chuyuan, et al.
Published: (2025)
by: Li, Chuyuan, et al.
Published: (2025)
Infusing Theory of Mind into Socially Intelligent LLM Agents
by: Hwang, EunJeong, et al.
Published: (2025)
by: Hwang, EunJeong, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
by: González-Pizarro, Felipe, et al.
Published: (2024)
by: González-Pizarro, Felipe, et al.
Published: (2024)
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025)
by: Boisvert, Léo, et al.
Published: (2025)
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
by: Boisvert, Léo, et al.
Published: (2024)
by: Boisvert, Léo, et al.
Published: (2024)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
by: Rodriguez, Juan A., et al.
Published: (2023)
by: Rodriguez, Juan A., et al.
Published: (2023)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
Similar Items
-
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024) -
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
by: Abaskohi, Amirhossein, et al.
Published: (2025) -
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026) -
BlockLLM: Memory-Efficient Adaptation of LLMs by Selecting and Optimizing the Right Coordinate Blocks
by: Ramesh, Amrutha Varshini, et al.
Published: (2024) -
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
by: Ashok, Arjun, et al.
Published: (2023)