Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
Fuente:
arXiv
Saved in:
| Main Authors: | Sundararajan, Barkavi, Sripada, Somayajulu, Reiter, Ehud |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo
by: Sundararajan, Barkavi, et al.
Published: (2024)
by: Sundararajan, Barkavi, et al.
Published: (2024)
We Should Evaluate Real-World Impact
by: Reiter, Ehud
Published: (2025)
by: Reiter, Ehud
Published: (2025)
NLG Evaluation: Past, Present, Future
by: Reiter, Ehud
Published: (2026)
by: Reiter, Ehud
Published: (2026)
Natural Language Generation
by: Reiter, Ehud
Published: (2025)
by: Reiter, Ehud
Published: (2025)
Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2025)
by: Bhatia, Gagan, et al.
Published: (2025)
Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative Analysis
by: Babakov, Nikolay, et al.
Published: (2024)
by: Babakov, Nikolay, et al.
Published: (2024)
Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models
by: Sivaprasad, Adarsa, et al.
Published: (2024)
by: Sivaprasad, Adarsa, et al.
Published: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
QFMTS: Generating Query-Focused Summaries over Multi-Table Inputs
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
Evaluating Language Translation Models by Playing Telephone
by: Saba, Syeda Jannatus, et al.
Published: (2025)
by: Saba, Syeda Jannatus, et al.
Published: (2025)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
by: Cheng, Letian, et al.
Published: (2026)
by: Cheng, Letian, et al.
Published: (2026)
From Role-Play to Drama-Interaction: An LLM Solution
by: Wu, Weiqi, et al.
Published: (2024)
by: Wu, Weiqi, et al.
Published: (2024)
Can LLM Teams Play What? Where? When?
by: Kotelnikova, Anastasia, et al.
Published: (2026)
by: Kotelnikova, Anastasia, et al.
Published: (2026)
PsyPlay: Personality-Infused Role-Playing Conversational Agents
by: Yang, Tao, et al.
Published: (2025)
by: Yang, Tao, et al.
Published: (2025)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
by: Struppek, Lukas, et al.
Published: (2025)
by: Struppek, Lukas, et al.
Published: (2025)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
by: Lu, Dongxu, et al.
Published: (2025)
by: Lu, Dongxu, et al.
Published: (2025)
DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents
by: Chuang, Yun-Shiuan, et al.
Published: (2025)
by: Chuang, Yun-Shiuan, et al.
Published: (2025)
Play to Generalize: Learning to Reason Through Game Play
by: Xie, Yunfei, et al.
Published: (2025)
by: Xie, Yunfei, et al.
Published: (2025)
Role-Playing Evaluation for Large Language Models
by: Boudouri, Yassine El, et al.
Published: (2025)
by: Boudouri, Yassine El, et al.
Published: (2025)
SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
by: Su, Xuanbo, et al.
Published: (2026)
by: Su, Xuanbo, et al.
Published: (2026)
Better LLM Reasoning via Dual-Play
by: Zhang, Zhengxin, et al.
Published: (2025)
by: Zhang, Zhengxin, et al.
Published: (2025)
Input Order Shapes LLM Semantic Alignment in Multi-Document Summarization
by: Ma, Jing
Published: (2025)
by: Ma, Jing
Published: (2025)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
by: Proietti, Michela, et al.
Published: (2025)
by: Proietti, Michela, et al.
Published: (2025)
Towards Sneaking as a Playful Input Modality for Virtual Environments
by: Cmentowski, Sebastian, et al.
Published: (2021)
by: Cmentowski, Sebastian, et al.
Published: (2021)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization
by: Shan, Baocai, et al.
Published: (2026)
by: Shan, Baocai, et al.
Published: (2026)
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
by: Singh, Punit Kumar, et al.
Published: (2025)
by: Singh, Punit Kumar, et al.
Published: (2025)
Textual Summarisation of Large Sets: Towards a General Approach
by: Kuptavanich, Kittipitch, et al.
Published: (2024)
by: Kuptavanich, Kittipitch, et al.
Published: (2024)
Codifying Character Logic in Role-Playing
by: Peng, Letian, et al.
Published: (2025)
by: Peng, Letian, et al.
Published: (2025)
Harnessing the Plug-and-Play Controller by Prompting
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
by: Zhang, Wenyuan, et al.
Published: (2024)
by: Zhang, Wenyuan, et al.
Published: (2024)
Learning to Play Like Humans: A Framework for LLM Adaptation in Interactive Fiction Games
by: Zhang, Jinming, et al.
Published: (2025)
by: Zhang, Jinming, et al.
Published: (2025)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
by: Gero, Zelalem, et al.
Published: (2024)
by: Gero, Zelalem, et al.
Published: (2024)
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
by: Tu, Quan, et al.
Published: (2024)
by: Tu, Quan, et al.
Published: (2024)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
by: Djiré, Albérick Euraste, et al.
Published: (2025)
by: Djiré, Albérick Euraste, et al.
Published: (2025)
Similar Items
-
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo
by: Sundararajan, Barkavi, et al.
Published: (2024) -
We Should Evaluate Real-World Impact
by: Reiter, Ehud
Published: (2025) -
NLG Evaluation: Past, Present, Future
by: Reiter, Ehud
Published: (2026) -
Natural Language Generation
by: Reiter, Ehud
Published: (2025) -
Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2025)