Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
Fuente:
arXiv
Salvato in:
| Autori principali: | Mondal, Ishani, Bhardwaj, Shweta |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
di: Lu, Janna
Pubblicazione: (2025)
di: Lu, Janna
Pubblicazione: (2025)
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
di: Liang, Yijun, et al.
Pubblicazione: (2025)
di: Liang, Yijun, et al.
Pubblicazione: (2025)
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
di: Schlegel, Udo, et al.
Pubblicazione: (2025)
di: Schlegel, Udo, et al.
Pubblicazione: (2025)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024)
di: Hou, Yufang, et al.
Pubblicazione: (2024)
BEExAI: Benchmark to Evaluate Explainable AI
di: Sithakoul, Samuel, et al.
Pubblicazione: (2024)
di: Sithakoul, Samuel, et al.
Pubblicazione: (2024)
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
di: Wu, Xiaobao, et al.
Pubblicazione: (2024)
di: Wu, Xiaobao, et al.
Pubblicazione: (2024)
K-QA: A Real-World Medical Q&A Benchmark
di: Manes, Itay, et al.
Pubblicazione: (2024)
di: Manes, Itay, et al.
Pubblicazione: (2024)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
di: Gupta, Rushil, et al.
Pubblicazione: (2025)
di: Gupta, Rushil, et al.
Pubblicazione: (2025)
We Should Chart an Atlas of All the World's Models
di: Horwitz, Eliahu, et al.
Pubblicazione: (2025)
di: Horwitz, Eliahu, et al.
Pubblicazione: (2025)
XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
di: Xiao, Yuzhuo, et al.
Pubblicazione: (2025)
di: Xiao, Yuzhuo, et al.
Pubblicazione: (2025)
Evaluating the Robustness and Accuracy of Text Watermarking Under Real-World Cross-Lingual Manipulations
di: Ghanim, Mansour Al, et al.
Pubblicazione: (2025)
di: Ghanim, Mansour Al, et al.
Pubblicazione: (2025)
RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
di: Jin, Zhuoran, et al.
Pubblicazione: (2024)
di: Jin, Zhuoran, et al.
Pubblicazione: (2024)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
di: Ling Team, et al.
Pubblicazione: (2025)
di: Ling Team, et al.
Pubblicazione: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
di: Long, Xiang, et al.
Pubblicazione: (2026)
di: Long, Xiang, et al.
Pubblicazione: (2026)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
di: Liu, Yixin, et al.
Pubblicazione: (2023)
di: Liu, Yixin, et al.
Pubblicazione: (2023)
Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Romanized Scripts in a Real World Setting
di: Khullar, Manurag, et al.
Pubblicazione: (2025)
di: Khullar, Manurag, et al.
Pubblicazione: (2025)
OWLViz: An Open-World Benchmark for Visual Question Answering
di: Nguyen, Thuy, et al.
Pubblicazione: (2025)
di: Nguyen, Thuy, et al.
Pubblicazione: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
di: Ren, Richard, et al.
Pubblicazione: (2024)
di: Ren, Richard, et al.
Pubblicazione: (2024)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
di: Wang, Jianghui, et al.
Pubblicazione: (2025)
di: Wang, Jianghui, et al.
Pubblicazione: (2025)
How much reliable is ChatGPT's prediction on Information Extraction under Input Perturbations?
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
di: Dai, Lu, et al.
Pubblicazione: (2025)
di: Dai, Lu, et al.
Pubblicazione: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
Privacy Evaluation Benchmarks for NLP Models
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
RelevAI-Reviewer: A Benchmark on AI Reviewers for Survey Paper Relevance
di: Couto, Paulo Henrique, et al.
Pubblicazione: (2024)
di: Couto, Paulo Henrique, et al.
Pubblicazione: (2024)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
di: Ye, Suyu, et al.
Pubblicazione: (2025)
di: Ye, Suyu, et al.
Pubblicazione: (2025)
From Task Solving to Robust Real-World Adaptation in LLM Agents
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
di: Shaier, Sagi, et al.
Pubblicazione: (2023)
di: Shaier, Sagi, et al.
Pubblicazione: (2023)
A Generative AI Framework for Intelligent Utility Billing CO 2 Analytics and Sustainable Resource Optimisation
di: Manjunath, Pavan, et al.
Pubblicazione: (2026)
di: Manjunath, Pavan, et al.
Pubblicazione: (2026)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
di: Guo, Zikang, et al.
Pubblicazione: (2025)
di: Guo, Zikang, et al.
Pubblicazione: (2025)
Divide-or-Conquer? Which Part Should You Distill Your LLM?
di: Wu, Zhuofeng, et al.
Pubblicazione: (2024)
di: Wu, Zhuofeng, et al.
Pubblicazione: (2024)
Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models
di: Islam, Mohammed Saidul, et al.
Pubblicazione: (2026)
di: Islam, Mohammed Saidul, et al.
Pubblicazione: (2026)
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
di: Huet, Alexis, et al.
Pubblicazione: (2025)
di: Huet, Alexis, et al.
Pubblicazione: (2025)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
di: Zhang, Wenbo, et al.
Pubblicazione: (2025)
di: Zhang, Wenbo, et al.
Pubblicazione: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
di: Lu, Janna
Pubblicazione: (2025) -
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
di: Zhong, Shanshan, et al.
Pubblicazione: (2026) -
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
di: Liang, Yijun, et al.
Pubblicazione: (2025) -
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
di: Schlegel, Udo, et al.
Pubblicazione: (2025) -
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024)