Benchmarking LLMs on the Semantic Overlap Summarization Task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Salvador, John, Bansal, Naman, Akter, Mousumi, Sarkar, Souvika, Das, Anupam, Karmaker, Shubhra Kanti |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FaNS: a Facet-based Narrative Similarity Metric
von: Akter, Mousumi, et al.
Veröffentlicht: (2023)
von: Akter, Mousumi, et al.
Veröffentlicht: (2023)
LLMs as On-demand Customizable Service
von: Sarkar, Souvika, et al.
Veröffentlicht: (2024)
von: Sarkar, Souvika, et al.
Veröffentlicht: (2024)
Redundancy Aware Multi-Reference Based Gainwise Evaluation of Extractive Summarization
von: Akter, Mousumi, et al.
Veröffentlicht: (2023)
von: Akter, Mousumi, et al.
Veröffentlicht: (2023)
Processing Natural Language on Embedded Devices: How Well Do Transformer Models Perform?
von: Sarkar, Souvika, et al.
Veröffentlicht: (2023)
von: Sarkar, Souvika, et al.
Veröffentlicht: (2023)
Revisiting Word Embeddings in the LLM Era
von: Mahajan, Yash, et al.
Veröffentlicht: (2024)
von: Mahajan, Yash, et al.
Veröffentlicht: (2024)
LLMs as Meta-Reviewers' Assistants: A Case Study
von: Hossain, Eftekhar, et al.
Veröffentlicht: (2024)
von: Hossain, Eftekhar, et al.
Veröffentlicht: (2024)
Zero-Shot Multi-Label Classification of Bangla Documents: Large Decoders Vs. Classic Encoders
von: Sarkar, Souvika, et al.
Veröffentlicht: (2025)
von: Sarkar, Souvika, et al.
Veröffentlicht: (2025)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2025)
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2025)
Set-Theoretic Compositionality of Sentence Embeddings
von: Bansal, Naman, et al.
Veröffentlicht: (2025)
von: Bansal, Naman, et al.
Veröffentlicht: (2025)
SNNLP: Energy-Efficient Natural Language Processing Using Spiking Neural Networks
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2024)
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2024)
Pitfalls of Evaluating Language Models with Open Benchmarks
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
Are LLMs Ready to Replace Bangla Annotators?
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2026)
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2026)
A Comprehensive Survey on Legal Summarization: Challenges and Future Directions
von: Akter, Mousumi, et al.
Veröffentlicht: (2025)
von: Akter, Mousumi, et al.
Veröffentlicht: (2025)
Large Language Models for IT Automation Tasks: Are We There Yet?
von: Hassan, Md Mahadi, et al.
Veröffentlicht: (2025)
von: Hassan, Md Mahadi, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on Urdu Idiom Translation
von: Khan, Muhammad Farmal, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Farmal, et al.
Veröffentlicht: (2025)
SimPal: Towards a Meta-Conversational Framework to Understand Teacher's Instructional Goals for K-12 Physics
von: Farhana, Effat, et al.
Veröffentlicht: (2024)
von: Farhana, Effat, et al.
Veröffentlicht: (2024)
Moneyball with LLMs: Analyzing Tabular Summarization in Sports Narratives
von: Upadhyay, Ritam, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ritam, et al.
Veröffentlicht: (2025)
Automatic Summarization of Long Documents
von: Chhibbar, Naman, et al.
Veröffentlicht: (2024)
von: Chhibbar, Naman, et al.
Veröffentlicht: (2024)
CREST: Universal Safety Guardrails Through Cluster-Guided Cross-Lingual Transfer
von: Bansal, Lavish, et al.
Veröffentlicht: (2025)
von: Bansal, Lavish, et al.
Veröffentlicht: (2025)
Sarcasm Detection on Reddit Using Classical Machine Learning and Feature Engineering
von: Karmaker, Subrata
Veröffentlicht: (2025)
von: Karmaker, Subrata
Veröffentlicht: (2025)
LexSumm and LexT5: Benchmarking and Modeling Legal Summarization Tasks in English
von: Santosh, T. Y. S. S., et al.
Veröffentlicht: (2024)
von: Santosh, T. Y. S. S., et al.
Veröffentlicht: (2024)
On Positional Bias of Faithfulness for Long-form Summarization
von: Wan, David, et al.
Veröffentlicht: (2024)
von: Wan, David, et al.
Veröffentlicht: (2024)
An Evaluation Benchmark for Autoformalization in Lean4
von: Gulati, Aryan, et al.
Veröffentlicht: (2024)
von: Gulati, Aryan, et al.
Veröffentlicht: (2024)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024)
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024)
The Bias is in the Details: An Assessment of Cognitive Bias in LLMs
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2025)
von: Knipper, R. Alexander, et al.
Veröffentlicht: (2025)
Evaluating LLMs and Pre-trained Models for Text Summarization Across Diverse Datasets
von: Rehman, Tohida, et al.
Veröffentlicht: (2025)
von: Rehman, Tohida, et al.
Veröffentlicht: (2025)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
Investigating Hallucination in Conversations for Low Resource Languages
von: Das, Amit, et al.
Veröffentlicht: (2025)
von: Das, Amit, et al.
Veröffentlicht: (2025)
Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based Method
von: Vipulanandan, Pragatheeswaran, et al.
Veröffentlicht: (2026)
von: Vipulanandan, Pragatheeswaran, et al.
Veröffentlicht: (2026)
Don't Judge a Book by its Cover: Testing LLMs' Robustness Under Logical Obfuscation
von: Borah, Abhilekh, et al.
Veröffentlicht: (2026)
von: Borah, Abhilekh, et al.
Veröffentlicht: (2026)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
von: Atil, Berk, et al.
Veröffentlicht: (2025)
von: Atil, Berk, et al.
Veröffentlicht: (2025)
Mitigating Semantic Drift: Evaluating LLMs' Efficacy in Psychotherapy through MI Dialogue Summarization
von: Kumar, Vivek, et al.
Veröffentlicht: (2025)
von: Kumar, Vivek, et al.
Veröffentlicht: (2025)
Enriching and Controlling Global Semantics for Text Summarization
von: Nguyen, Thong, et al.
Veröffentlicht: (2021)
von: Nguyen, Thong, et al.
Veröffentlicht: (2021)
Demystifying Legalese: An Automated Approach for Summarizing and Analyzing Overlaps in Privacy Policies and Terms of Service
von: Soneji, Shikha, et al.
Veröffentlicht: (2024)
von: Soneji, Shikha, et al.
Veröffentlicht: (2024)
Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization
von: Mei, Xiaoyong, et al.
Veröffentlicht: (2026)
von: Mei, Xiaoyong, et al.
Veröffentlicht: (2026)
Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance
von: Chung, Woojin, et al.
Veröffentlicht: (2026)
von: Chung, Woojin, et al.
Veröffentlicht: (2026)
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
von: Varshney, Neeraj, et al.
Veröffentlicht: (2024)
von: Varshney, Neeraj, et al.
Veröffentlicht: (2024)
OrderSum: Semantic Sentence Ordering for Extractive Summarization
von: Kwon, Taewan, et al.
Veröffentlicht: (2025)
von: Kwon, Taewan, et al.
Veröffentlicht: (2025)
Evaluating the Ability of LLMs to Solve Semantics-Aware Process Mining Tasks
von: Rebmann, Adrian, et al.
Veröffentlicht: (2024)
von: Rebmann, Adrian, et al.
Veröffentlicht: (2024)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FaNS: a Facet-based Narrative Similarity Metric
von: Akter, Mousumi, et al.
Veröffentlicht: (2023) -
LLMs as On-demand Customizable Service
von: Sarkar, Souvika, et al.
Veröffentlicht: (2024) -
Redundancy Aware Multi-Reference Based Gainwise Evaluation of Extractive Summarization
von: Akter, Mousumi, et al.
Veröffentlicht: (2023) -
Processing Natural Language on Embedded Devices: How Well Do Transformer Models Perform?
von: Sarkar, Souvika, et al.
Veröffentlicht: (2023) -
Revisiting Word Embeddings in the LLM Era
von: Mahajan, Yash, et al.
Veröffentlicht: (2024)