Multi-Dimensional Evaluation of Text Summarization with In-Context Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Jain, Sameer, Keshava, Vaishakh, Sathyendra, Swarnashree Mysore, Fernandes, Patrick, Liu, Pengfei, Neubig, Graham, Zhou, Chunting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models
di: Xie, Tian, et al.
Pubblicazione: (2025)
di: Xie, Tian, et al.
Pubblicazione: (2025)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
di: Chern, Steffi, et al.
Pubblicazione: (2024)
di: Chern, Steffi, et al.
Pubblicazione: (2024)
Causal Understanding For Video Question Answering
di: Guda, Bhanu Prakash Reddy, et al.
Pubblicazione: (2024)
di: Guda, Bhanu Prakash Reddy, et al.
Pubblicazione: (2024)
Reinforcement Learning with Backtracking Feedback
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
di: Song, Woomin, et al.
Pubblicazione: (2025)
di: Song, Woomin, et al.
Pubblicazione: (2025)
Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning
di: Ryu, Sangwon, et al.
Pubblicazione: (2024)
di: Ryu, Sangwon, et al.
Pubblicazione: (2024)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
di: Fernandes, Patrick, et al.
Pubblicazione: (2025)
di: Fernandes, Patrick, et al.
Pubblicazione: (2025)
An Incomplete Loop: Instruction Inference, Instruction Following, and In-context Learning in Language Models
di: Liu, Emmy, et al.
Pubblicazione: (2024)
di: Liu, Emmy, et al.
Pubblicazione: (2024)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
di: Wang, Zizhao, et al.
Pubblicazione: (2025)
di: Wang, Zizhao, et al.
Pubblicazione: (2025)
Backtracking for Safety
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
di: Bertsch, Amanda, et al.
Pubblicazione: (2025)
di: Bertsch, Amanda, et al.
Pubblicazione: (2025)
Multi-LLM Text Summarization
di: Fang, Jiangnan, et al.
Pubblicazione: (2024)
di: Fang, Jiangnan, et al.
Pubblicazione: (2024)
In-Context Learning with Long-Context Models: An In-Depth Exploration
di: Bertsch, Amanda, et al.
Pubblicazione: (2024)
di: Bertsch, Amanda, et al.
Pubblicazione: (2024)
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
di: Xiao, Emily, et al.
Pubblicazione: (2025)
di: Xiao, Emily, et al.
Pubblicazione: (2025)
Go-Browse: Training Web Agents with Structured Exploration
di: Gandhi, Apurva, et al.
Pubblicazione: (2025)
di: Gandhi, Apurva, et al.
Pubblicazione: (2025)
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
di: Tjuatja, Lindia, et al.
Pubblicazione: (2025)
di: Tjuatja, Lindia, et al.
Pubblicazione: (2025)
An Empirical Comparison of Text Summarization: A Multi-Dimensional Evaluation of Large Language Models
di: Janakiraman, Anantharaman, et al.
Pubblicazione: (2025)
di: Janakiraman, Anantharaman, et al.
Pubblicazione: (2025)
Alignment for Honesty
di: Yang, Yuqing, et al.
Pubblicazione: (2023)
di: Yang, Yuqing, et al.
Pubblicazione: (2023)
Midtraining Bridges Pretraining and Posttraining Distributions
di: Liu, Emmy, et al.
Pubblicazione: (2025)
di: Liu, Emmy, et al.
Pubblicazione: (2025)
Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization
di: Sun, Shichao, et al.
Pubblicazione: (2024)
di: Sun, Shichao, et al.
Pubblicazione: (2024)
Effective Strategies for Asynchronous Software Engineering Agents
di: Geng, Jiayi, et al.
Pubblicazione: (2026)
di: Geng, Jiayi, et al.
Pubblicazione: (2026)
Better Instruction-Following Through Minimum Bayes Risk
di: Wu, Ian, et al.
Pubblicazione: (2024)
di: Wu, Ian, et al.
Pubblicazione: (2024)
Towards Automatic Evaluation for Image Transcreation
di: Khanuja, Simran, et al.
Pubblicazione: (2024)
di: Khanuja, Simran, et al.
Pubblicazione: (2024)
On the Role of Summary Content Units in Text Summarization Evaluation
di: Nawrath, Marcel, et al.
Pubblicazione: (2024)
di: Nawrath, Marcel, et al.
Pubblicazione: (2024)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
di: Kaneko, Masahiro, et al.
Pubblicazione: (2023)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2023)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
di: Song, Yueqi, et al.
Pubblicazione: (2024)
di: Song, Yueqi, et al.
Pubblicazione: (2024)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
di: Zhang, Charlie, et al.
Pubblicazione: (2025)
di: Zhang, Charlie, et al.
Pubblicazione: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
di: Lin, Zhiqiu, et al.
Pubblicazione: (2024)
di: Lin, Zhiqiu, et al.
Pubblicazione: (2024)
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches
di: Pratapa, Adithya, et al.
Pubblicazione: (2025)
di: Pratapa, Adithya, et al.
Pubblicazione: (2025)
On Learning to Summarize with Large Language Models as References
di: Liu, Yixin, et al.
Pubblicazione: (2023)
di: Liu, Yixin, et al.
Pubblicazione: (2023)
Instruction-tuned Language Models are Better Knowledge Learners
di: Jiang, Zhengbao, et al.
Pubblicazione: (2024)
di: Jiang, Zhengbao, et al.
Pubblicazione: (2024)
Beyond Browsing: API-Based Web Agents
di: Song, Yueqi, et al.
Pubblicazione: (2024)
di: Song, Yueqi, et al.
Pubblicazione: (2024)
Is Context Helpful for Chat Translation Evaluation?
di: Agrawal, Sweta, et al.
Pubblicazione: (2024)
di: Agrawal, Sweta, et al.
Pubblicazione: (2024)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
di: Yayavaram, Arnav, et al.
Pubblicazione: (2025)
di: Yayavaram, Arnav, et al.
Pubblicazione: (2025)
DTELS: Towards Dynamic Granularity of Timeline Summarization
di: Zhang, Chenlong, et al.
Pubblicazione: (2024)
di: Zhang, Chenlong, et al.
Pubblicazione: (2024)
DUAL: Diversity and Uncertainty Active Learning for Text Summarization
di: Giouroukis, Petros Stylianos, et al.
Pubblicazione: (2025)
di: Giouroukis, Petros Stylianos, et al.
Pubblicazione: (2025)
Accumulating Context Changes the Beliefs of Language Models
di: Geng, Jiayi, et al.
Pubblicazione: (2025)
di: Geng, Jiayi, et al.
Pubblicazione: (2025)
Data-efficient Performance Modeling via Pre-training
di: Liu, Chunting, et al.
Pubblicazione: (2025)
di: Liu, Chunting, et al.
Pubblicazione: (2025)
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
di: Song, Yueqi, et al.
Pubblicazione: (2025)
di: Song, Yueqi, et al.
Pubblicazione: (2025)
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text
di: Ginn, Michael, et al.
Pubblicazione: (2024)
di: Ginn, Michael, et al.
Pubblicazione: (2024)
Documenti analoghi
-
BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models
di: Xie, Tian, et al.
Pubblicazione: (2025) -
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
di: Chern, Steffi, et al.
Pubblicazione: (2024) -
Causal Understanding For Video Question Answering
di: Guda, Bhanu Prakash Reddy, et al.
Pubblicazione: (2024) -
Reinforcement Learning with Backtracking Feedback
di: Sel, Bilgehan, et al.
Pubblicazione: (2026) -
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
di: Song, Woomin, et al.
Pubblicazione: (2025)