InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Sahu, Gaurav, Puri, Abhay, Rodriguez, Juan, Abaskohi, Amirhossein, Chegini, Mohammad, Drouin, Alexandre, Taslakian, Perouz, Zantedeschi, Valentina, Lacoste, Alexandre, Vazquez, David, Chapados, Nicolas, Pal, Christopher, Mudumba, Sai Rajeswar, Laradji, Issam Hadj |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
Learning to Defer for Causal Discovery with Imperfect Experts
by: Clivio, Oscar, et al.
Published: (2025)
by: Clivio, Oscar, et al.
Published: (2025)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
by: Ashok, Arjun, et al.
Published: (2023)
by: Ashok, Arjun, et al.
Published: (2023)
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
by: Rodriguez, Juan A., et al.
Published: (2023)
by: Rodriguez, Juan A., et al.
Published: (2023)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization
by: Sahu, Gaurav, et al.
Published: (2023)
by: Sahu, Gaurav, et al.
Published: (2023)
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025)
by: Boisvert, Léo, et al.
Published: (2025)
DRBench: A Realistic Benchmark for Enterprise Deep Research
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
by: Drouin, Alexandre, et al.
Published: (2024)
by: Drouin, Alexandre, et al.
Published: (2024)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Context is Key: A Benchmark for Forecasting with Essential Textual Information
by: Williams, Andrew Robert, et al.
Published: (2024)
by: Williams, Andrew Robert, et al.
Published: (2024)
Dr-CiK: A Testbed for Foresight-Driven Agents
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
Choreographer: Learning and Adapting Skills in Imagination
by: Mazzaglia, Pietro, et al.
Published: (2022)
by: Mazzaglia, Pietro, et al.
Published: (2022)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
Hierarchical Retrieval at Scale: Bridging Transparency and Efficiency
by: Gupta, Shubham, et al.
Published: (2025)
by: Gupta, Shubham, et al.
Published: (2025)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
by: Ashok, Arjun, et al.
Published: (2025)
by: Ashok, Arjun, et al.
Published: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
by: Boisvert, Léo, et al.
Published: (2024)
by: Boisvert, Léo, et al.
Published: (2024)
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
by: Boisvert, Leo, et al.
Published: (2025)
by: Boisvert, Leo, et al.
Published: (2025)
Seventeen 2 Micron All Sky Survey (2MASS) hypervelocity stars (HVS) from Gaia DR3
by: Mudumba, Parthasarathy
Published: (2024)
by: Mudumba, Parthasarathy
Published: (2024)
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
by: Sadeghi, Pouya, et al.
Published: (2024)
by: Sadeghi, Pouya, et al.
Published: (2024)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
by: Tsirigotis, Christos, et al.
Published: (2025)
by: Tsirigotis, Christos, et al.
Published: (2025)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
by: Rodriguez, Juan, et al.
Published: (2024)
by: Rodriguez, Juan, et al.
Published: (2024)
SSLR: A Semi-Supervised Learning Method for Isolated Sign Language Recognition
by: Algafri, Hasan, et al.
Published: (2025)
by: Algafri, Hasan, et al.
Published: (2025)
Similar Items
-
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026) -
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024) -
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025) -
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024) -
Learning to Defer for Causal Discovery with Imperfect Experts
by: Clivio, Oscar, et al.
Published: (2025)