Benchmarking Table Comprehension In The Wild
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Yikang, Zhu, Yi, Xie, Rand, Liu, Yizhi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
di: Wang, Pengyu, et al.
Pubblicazione: (2026)
di: Wang, Pengyu, et al.
Pubblicazione: (2026)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
di: Wu, Xianjie, et al.
Pubblicazione: (2024)
di: Wu, Xianjie, et al.
Pubblicazione: (2024)
Benchmarking LLM Tool-Use in the Wild
di: Yu, Peijie, et al.
Pubblicazione: (2026)
di: Yu, Peijie, et al.
Pubblicazione: (2026)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
di: Zhu, Zhiying, et al.
Pubblicazione: (2024)
di: Zhu, Zhiying, et al.
Pubblicazione: (2024)
Membership Inference on LLMs in the Wild
di: Yi, Jiatong, et al.
Pubblicazione: (2026)
di: Yi, Jiatong, et al.
Pubblicazione: (2026)
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
di: Wu, Pengzuo, et al.
Pubblicazione: (2025)
di: Wu, Pengzuo, et al.
Pubblicazione: (2025)
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
di: Pan, Changzai, et al.
Pubblicazione: (2026)
di: Pan, Changzai, et al.
Pubblicazione: (2026)
PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models
di: Wang, Yuwen, et al.
Pubblicazione: (2026)
di: Wang, Yuwen, et al.
Pubblicazione: (2026)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
di: Ding, Shuangrui, et al.
Pubblicazione: (2026)
di: Ding, Shuangrui, et al.
Pubblicazione: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
di: Yan, Lian, et al.
Pubblicazione: (2025)
di: Yan, Lian, et al.
Pubblicazione: (2025)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
di: Mundada, Gagan, et al.
Pubblicazione: (2025)
di: Mundada, Gagan, et al.
Pubblicazione: (2025)
OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
di: Zhang, Xinyu, et al.
Pubblicazione: (2026)
di: Zhang, Xinyu, et al.
Pubblicazione: (2026)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
di: Wu, Siyang, et al.
Pubblicazione: (2025)
di: Wu, Siyang, et al.
Pubblicazione: (2025)
CRAG -- Comprehensive RAG Benchmark
di: Yang, Xiao, et al.
Pubblicazione: (2024)
di: Yang, Xiao, et al.
Pubblicazione: (2024)
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
di: Wei, Kaiwen, et al.
Pubblicazione: (2025)
di: Wei, Kaiwen, et al.
Pubblicazione: (2025)
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
di: Liu, Yujian, et al.
Pubblicazione: (2026)
di: Liu, Yujian, et al.
Pubblicazione: (2026)
FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2024)
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2024)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
di: Zhang, Tao, et al.
Pubblicazione: (2024)
di: Zhang, Tao, et al.
Pubblicazione: (2024)
COOL: Comprehensive Knowledge Enhanced Prompt Learning for Domain Adaptive Few-shot Fake News Detection
di: Ouyang, Yi, et al.
Pubblicazione: (2024)
di: Ouyang, Yi, et al.
Pubblicazione: (2024)
Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
di: Huang, Shijue, et al.
Pubblicazione: (2024)
di: Huang, Shijue, et al.
Pubblicazione: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
di: Li, Ce, et al.
Pubblicazione: (2025)
di: Li, Ce, et al.
Pubblicazione: (2025)
The Mask of Civility: Benchmarking Chinese Mock Politeness Comprehension in Large Language Models
di: Zhang, Yitong, et al.
Pubblicazione: (2026)
di: Zhang, Yitong, et al.
Pubblicazione: (2026)
AKEW: Assessing Knowledge Editing in the Wild
di: Wu, Xiaobao, et al.
Pubblicazione: (2024)
di: Wu, Xiaobao, et al.
Pubblicazione: (2024)
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
di: Inc, Xiaohongshu
Pubblicazione: (2026)
di: Inc, Xiaohongshu
Pubblicazione: (2026)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
di: Tan, Rongyuan, et al.
Pubblicazione: (2026)
di: Tan, Rongyuan, et al.
Pubblicazione: (2026)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
di: Wang, Yumeng, et al.
Pubblicazione: (2025)
di: Wang, Yumeng, et al.
Pubblicazione: (2025)
YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
di: Yu, Deshui, et al.
Pubblicazione: (2025)
di: Yu, Deshui, et al.
Pubblicazione: (2025)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
GuessBench: Sensemaking Multimodal Creativity in the Wild
di: Zhu, Zifeng, et al.
Pubblicazione: (2025)
di: Zhu, Zifeng, et al.
Pubblicazione: (2025)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
di: Liu, Zhihang, et al.
Pubblicazione: (2025)
di: Liu, Zhihang, et al.
Pubblicazione: (2025)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
di: Yi, Jingwei, et al.
Pubblicazione: (2023)
di: Yi, Jingwei, et al.
Pubblicazione: (2023)
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering
di: Zhang, Xuanliang, et al.
Pubblicazione: (2025)
di: Zhang, Xuanliang, et al.
Pubblicazione: (2025)
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
di: Chen, Shisong, et al.
Pubblicazione: (2025)
di: Chen, Shisong, et al.
Pubblicazione: (2025)
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
Translation in the Wild
di: Balashov, Yuri
Pubblicazione: (2025)
di: Balashov, Yuri
Pubblicazione: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
di: Du, Mingxuan, et al.
Pubblicazione: (2025)
di: Du, Mingxuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
di: Wang, Pengyu, et al.
Pubblicazione: (2026) -
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
di: Zhang, Linhao, et al.
Pubblicazione: (2025) -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
di: Wu, Xianjie, et al.
Pubblicazione: (2024) -
Benchmarking LLM Tool-Use in the Wild
di: Yu, Peijie, et al.
Pubblicazione: (2026) -
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
di: Zhu, Zhiying, et al.
Pubblicazione: (2024)