We Should Evaluate Real-World Impact
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Reiter, Ehud |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NLG Evaluation: Past, Present, Future
von: Reiter, Ehud
Veröffentlicht: (2026)
von: Reiter, Ehud
Veröffentlicht: (2026)
Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
von: Sundararajan, Barkavi, et al.
Veröffentlicht: (2025)
von: Sundararajan, Barkavi, et al.
Veröffentlicht: (2025)
Natural Language Generation
von: Reiter, Ehud
Veröffentlicht: (2025)
von: Reiter, Ehud
Veröffentlicht: (2025)
Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2024)
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2024)
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo
von: Sundararajan, Barkavi, et al.
Veröffentlicht: (2024)
von: Sundararajan, Barkavi, et al.
Veröffentlicht: (2024)
Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative Analysis
von: Babakov, Nikolay, et al.
Veröffentlicht: (2024)
von: Babakov, Nikolay, et al.
Veröffentlicht: (2024)
We Should Chart an Atlas of All the World's Models
von: Horwitz, Eliahu, et al.
Veröffentlicht: (2025)
von: Horwitz, Eliahu, et al.
Veröffentlicht: (2025)
Position: AI Evaluation Should Learn from How We Test Humans
von: Zhuang, Yan, et al.
Veröffentlicht: (2023)
von: Zhuang, Yan, et al.
Veröffentlicht: (2023)
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
von: Mondal, Ishani, et al.
Veröffentlicht: (2026)
von: Mondal, Ishani, et al.
Veröffentlicht: (2026)
Textual Summarisation of Large Sets: Towards a General Approach
von: Kuptavanich, Kittipitch, et al.
Veröffentlicht: (2024)
von: Kuptavanich, Kittipitch, et al.
Veröffentlicht: (2024)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
von: Li, Karen Jia-Hui, et al.
Veröffentlicht: (2025)
von: Li, Karen Jia-Hui, et al.
Veröffentlicht: (2025)
How Should We Model the Probability of a Language?
von: Dent, Rasul, et al.
Veröffentlicht: (2026)
von: Dent, Rasul, et al.
Veröffentlicht: (2026)
Just Because We Camp, Doesn't Mean We Should: The Ethics of Modelling Queer Voices
von: Sigurgeirsson, Atli, et al.
Veröffentlicht: (2024)
von: Sigurgeirsson, Atli, et al.
Veröffentlicht: (2024)
An End-to-End System for Culturally-Attuned Driving Feedback using a Dual-Component NLG Engine
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
Should We Still Pretrain Encoders with Masked Language Modeling?
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
Ask the experts: sourcing high-quality datasets for nutritional counselling through Human-AI collaboration
von: Balloccu, Simone, et al.
Veröffentlicht: (2024)
von: Balloccu, Simone, et al.
Veröffentlicht: (2024)
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
von: Alghisi, Simone, et al.
Veröffentlicht: (2024)
von: Alghisi, Simone, et al.
Veröffentlicht: (2024)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
Should We be Pedantic About Reasoning Errors in Machine Translation?
von: Bao, Calvin, et al.
Veröffentlicht: (2026)
von: Bao, Calvin, et al.
Veröffentlicht: (2026)
Are Word Embedding Methods Stable and Should We Care About It?
von: Borah, Angana, et al.
Veröffentlicht: (2021)
von: Borah, Angana, et al.
Veröffentlicht: (2021)
We Should Separate Memorization from Copyright
von: Haviv, Adi, et al.
Veröffentlicht: (2026)
von: Haviv, Adi, et al.
Veröffentlicht: (2026)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
von: Yin, Ziqi, et al.
Veröffentlicht: (2024)
von: Yin, Ziqi, et al.
Veröffentlicht: (2024)
Dense X Retrieval: What Retrieval Granularity Should We Use?
von: Chen, Tong, et al.
Veröffentlicht: (2023)
von: Chen, Tong, et al.
Veröffentlicht: (2023)
Real-World Summarization: When Evaluation Reaches Its Limits
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025)
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025)
Types for Grassroots Logic Programs
von: Shapiro, Ehud
Veröffentlicht: (2026)
von: Shapiro, Ehud
Veröffentlicht: (2026)
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
von: Barmina, Gianluca, et al.
Veröffentlicht: (2025)
von: Barmina, Gianluca, et al.
Veröffentlicht: (2025)
Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
von: Khoramfar, Ali, et al.
Veröffentlicht: (2025)
von: Khoramfar, Ali, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
Effectiveness of ChatGPT in explaining complex medical reports to patients
von: Sun, Mengxuan, et al.
Veröffentlicht: (2024)
von: Sun, Mengxuan, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Real-World Engineering Tasks
von: Heesch, Rene, et al.
Veröffentlicht: (2025)
von: Heesch, Rene, et al.
Veröffentlicht: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
von: Koto, Fajri
Veröffentlicht: (2024)
von: Koto, Fajri
Veröffentlicht: (2024)
TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles
von: Yu, Qingchen, et al.
Veröffentlicht: (2024)
von: Yu, Qingchen, et al.
Veröffentlicht: (2024)
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
von: Wagner, Eitan, et al.
Veröffentlicht: (2025)
von: Wagner, Eitan, et al.
Veröffentlicht: (2025)
Where Are We? Evaluating LLM Performance on African Languages
von: Adebara, Ife, et al.
Veröffentlicht: (2025)
von: Adebara, Ife, et al.
Veröffentlicht: (2025)
Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
von: Ochieng, Millicent, et al.
Veröffentlicht: (2024)
von: Ochieng, Millicent, et al.
Veröffentlicht: (2024)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
von: Wu, Yihao, et al.
Veröffentlicht: (2025)
von: Wu, Yihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NLG Evaluation: Past, Present, Future
von: Reiter, Ehud
Veröffentlicht: (2026) -
Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
von: Sundararajan, Barkavi, et al.
Veröffentlicht: (2025) -
Natural Language Generation
von: Reiter, Ehud
Veröffentlicht: (2025) -
Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2024) -
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo
von: Sundararajan, Barkavi, et al.
Veröffentlicht: (2024)