InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Zhenghao, Song, Yuanfeng, Chen, Xin, Liu, Chengzhong, Cui, Yakun, Cao, Caleb Chen, Han, Sirui, Guo, Yike |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
di: Zhu, Zhenghao, et al.
Pubblicazione: (2025)
di: Zhu, Zhenghao, et al.
Pubblicazione: (2025)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
di: Yakun, Cui, et al.
Pubblicazione: (2025)
di: Yakun, Cui, et al.
Pubblicazione: (2025)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
di: Yakun, Cui, et al.
Pubblicazione: (2026)
di: Yakun, Cui, et al.
Pubblicazione: (2026)
MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
di: Yakun, Cui, et al.
Pubblicazione: (2026)
di: Yakun, Cui, et al.
Pubblicazione: (2026)
DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning
di: Liu, Xiaochuan, et al.
Pubblicazione: (2025)
di: Liu, Xiaochuan, et al.
Pubblicazione: (2025)
Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI Collaboration
di: Han, Sirui, et al.
Pubblicazione: (2026)
di: Han, Sirui, et al.
Pubblicazione: (2026)
Automate Strategy Finding with LLM in Quant Investment
di: Kou, Zhizhuo, et al.
Pubblicazione: (2024)
di: Kou, Zhizhuo, et al.
Pubblicazione: (2024)
LLM Agent Swarm for Hypothesis-Driven Drug Discovery
di: Song, Kevin, et al.
Pubblicazione: (2025)
di: Song, Kevin, et al.
Pubblicazione: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
di: Li, Lujun, et al.
Pubblicazione: (2025)
di: Li, Lujun, et al.
Pubblicazione: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
di: Liu, Zhou, et al.
Pubblicazione: (2025)
di: Liu, Zhou, et al.
Pubblicazione: (2025)
An Electrocardiogram Multi-task Benchmark with Comprehensive Evaluations and Insightful Findings
di: Xu, Yuhao, et al.
Pubblicazione: (2025)
di: Xu, Yuhao, et al.
Pubblicazione: (2025)
Benchmarking Multi-National Value Alignment for Large Language Models
di: Shi, Weijie, et al.
Pubblicazione: (2025)
di: Shi, Weijie, et al.
Pubblicazione: (2025)
From AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM Agents
di: Xia, Sirui, et al.
Pubblicazione: (2026)
di: Xia, Sirui, et al.
Pubblicazione: (2026)
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong
di: Han, Sirui, et al.
Pubblicazione: (2025)
di: Han, Sirui, et al.
Pubblicazione: (2025)
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
di: Chen, Weiyi, et al.
Pubblicazione: (2026)
di: Chen, Weiyi, et al.
Pubblicazione: (2026)
Measuring Hong Kong Massive Multi-Task Language Understanding
di: Cao, Chuxue, et al.
Pubblicazione: (2025)
di: Cao, Chuxue, et al.
Pubblicazione: (2025)
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
di: Cao, Chuxue, et al.
Pubblicazione: (2025)
di: Cao, Chuxue, et al.
Pubblicazione: (2025)
SafeLawBench: Towards Safe Alignment of Large Language Models
di: Cao, Chuxue, et al.
Pubblicazione: (2025)
di: Cao, Chuxue, et al.
Pubblicazione: (2025)
LLM Unlearning Without an Expert Curated Dataset
di: Zhu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Zhu, Xiaoyuan, et al.
Pubblicazione: (2025)
Assessing Allele Frequency Information: A Study of Variant Curation Expert Panel Guidelines
di: Xiaoyan Wang, et al.
Pubblicazione: (2026)
di: Xiaoyan Wang, et al.
Pubblicazione: (2026)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
di: Zheng, Ruobing, et al.
Pubblicazione: (2026)
di: Zheng, Ruobing, et al.
Pubblicazione: (2026)
FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
di: Luo, Junyu, et al.
Pubblicazione: (2025)
di: Luo, Junyu, et al.
Pubblicazione: (2025)
Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent
di: Wu, Dongming, et al.
Pubblicazione: (2026)
di: Wu, Dongming, et al.
Pubblicazione: (2026)
ExpertQA: Expert-Curated Questions and Attributed Answers
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
AcademicEval: Live Long-Context LLM Benchmark
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
LLM-Driven Security Analysis for Cellular Networks: Vulnerability Discovery Agent and Benchmarking
di: Xie, Tian
Pubblicazione: (2026)
di: Xie, Tian
Pubblicazione: (2026)
Insight Agents: An LLM-Based Multi-Agent System for Data Insights
di: Bai, Jincheng, et al.
Pubblicazione: (2026)
di: Bai, Jincheng, et al.
Pubblicazione: (2026)
Interpretable AI-Driven Discovery of Terrain-Precipitation Relationships for Enhanced Climate Insights
di: Xu, Hao, et al.
Pubblicazione: (2023)
di: Xu, Hao, et al.
Pubblicazione: (2023)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
di: Liang, Yuanzhi, et al.
Pubblicazione: (2024)
di: Liang, Yuanzhi, et al.
Pubblicazione: (2024)
RW-TTT: Batched Serving for Request-Owned Test-Time Training State
di: Yang, Jian, et al.
Pubblicazione: (2026)
di: Yang, Jian, et al.
Pubblicazione: (2026)
Benchmarking Physics-Informed Time-Series Models for Operational Global Station Weather Forecasting
di: Han, Tao, et al.
Pubblicazione: (2024)
di: Han, Tao, et al.
Pubblicazione: (2024)
Optimizing Sequencing Coverage Depth in DNA Storage: Insights From DNA Storage Data
di: Cao, Ruiying, et al.
Pubblicazione: (2025)
di: Cao, Ruiying, et al.
Pubblicazione: (2025)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
di: Chen, Shengzhuang, et al.
Pubblicazione: (2025)
di: Chen, Shengzhuang, et al.
Pubblicazione: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
di: Tewolde, Emanuel, et al.
Pubblicazione: (2026)
di: Tewolde, Emanuel, et al.
Pubblicazione: (2026)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
di: Fish, Sara, et al.
Pubblicazione: (2025)
di: Fish, Sara, et al.
Pubblicazione: (2025)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Revolutionizing Bridge Operation and Maintenance with LLM-based Agents: An Overview of Applications and Insights
di: Chen, Xinyu, et al.
Pubblicazione: (2024)
di: Chen, Xinyu, et al.
Pubblicazione: (2024)
Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization
di: Liu, Shengchao, et al.
Pubblicazione: (2025)
di: Liu, Shengchao, et al.
Pubblicazione: (2025)
TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine
di: Cheng, Zihao, et al.
Pubblicazione: (2025)
di: Cheng, Zihao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
di: Zhu, Zhenghao, et al.
Pubblicazione: (2025) -
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
di: Yakun, Cui, et al.
Pubblicazione: (2025) -
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
di: Yakun, Cui, et al.
Pubblicazione: (2026) -
MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
di: Yakun, Cui, et al.
Pubblicazione: (2026) -
DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning
di: Liu, Xiaochuan, et al.
Pubblicazione: (2025)