AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Afzal, Anum, Chalumattu, Ribin, Matthes, Florian, Mascarell, Laura |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
par: Mascarell, Laura, et autres
Publié: (2024)
par: Mascarell, Laura, et autres
Publié: (2024)
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
par: Afzal, Anum, et autres
Publié: (2025)
par: Afzal, Anum, et autres
Publié: (2025)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
par: Afzal, Anum, et autres
Publié: (2025)
par: Afzal, Anum, et autres
Publié: (2025)
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
par: Afzal, Anum, et autres
Publié: (2025)
par: Afzal, Anum, et autres
Publié: (2025)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
par: Afzal, Anum, et autres
Publié: (2024)
par: Afzal, Anum, et autres
Publié: (2024)
Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
par: Afzal, Anum, et autres
Publié: (2025)
par: Afzal, Anum, et autres
Publié: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
par: Zhang, Tanghaoran, et autres
Publié: (2026)
par: Zhang, Tanghaoran, et autres
Publié: (2026)
Enhancing Answer Attribution for Faithful Text Generation with Large Language Models
par: Vladika, Juraj, et autres
Publié: (2024)
par: Vladika, Juraj, et autres
Publié: (2024)
Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting
par: Meisenbacher, Stephen, et autres
Publié: (2024)
par: Meisenbacher, Stephen, et autres
Publié: (2024)
A Comparative Analysis of Conversational Large Language Models in Knowledge-Based Text Generation
par: Schneider, Phillip, et autres
Publié: (2024)
par: Schneider, Phillip, et autres
Publié: (2024)
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
par: Van Veen, Dave, et autres
Publié: (2023)
par: Van Veen, Dave, et autres
Publié: (2023)
Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models
par: Vladika, Juraj, et autres
Publié: (2025)
par: Vladika, Juraj, et autres
Publié: (2025)
DP-MLM: Differentially Private Text Rewriting Using Masked Language Models
par: Meisenbacher, Stephen, et autres
Publié: (2024)
par: Meisenbacher, Stephen, et autres
Publié: (2024)
MedREQAL: Examining Medical Knowledge Recall of Large Language Models via Question Answering
par: Vladika, Juraj, et autres
Publié: (2024)
par: Vladika, Juraj, et autres
Publié: (2024)
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
par: Afzal, Anum, et autres
Publié: (2026)
par: Afzal, Anum, et autres
Publié: (2026)
With Privacy, Size Matters: On the Importance of Dataset Size in Differentially Private Text Rewriting
par: Meisenbacher, Stephen, et autres
Publié: (2025)
par: Meisenbacher, Stephen, et autres
Publié: (2025)
Just Rewrite It Again: A Post-Processing Method for Enhanced Semantic Similarity and Privacy Preservation of Differentially Private Rewritten Text
par: Meisenbacher, Stephen, et autres
Publié: (2024)
par: Meisenbacher, Stephen, et autres
Publié: (2024)
A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation
par: Meisenbacher, Stephen, et autres
Publié: (2026)
par: Meisenbacher, Stephen, et autres
Publié: (2026)
Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise: A Case Study on Chinese Legal Domain
par: wan, Zhen, et autres
Publié: (2023)
par: wan, Zhen, et autres
Publié: (2023)
Evaluating Large Language Models in Semantic Parsing for Conversational Question Answering over Knowledge Graphs
par: Schneider, Phillip, et autres
Publié: (2024)
par: Schneider, Phillip, et autres
Publié: (2024)
On the Impact of Noise in Differentially Private Text Rewriting
par: Meisenbacher, Stephen, et autres
Publié: (2025)
par: Meisenbacher, Stephen, et autres
Publié: (2025)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
par: Gu, Zhouhong, et autres
Publié: (2024)
par: Gu, Zhouhong, et autres
Publié: (2024)
NLP-KG: A System for Exploratory Search of Scientific Literature in Natural Language Processing
par: Schopf, Tim, et autres
Publié: (2024)
par: Schopf, Tim, et autres
Publié: (2024)
Scaling Up Summarization: Leveraging Large Language Models for Long Text Extractive Summarization
par: Hemamou, Léo, et autres
Publié: (2024)
par: Hemamou, Léo, et autres
Publié: (2024)
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
par: Guo, Xin, et autres
Publié: (2023)
par: Guo, Xin, et autres
Publié: (2023)
Comparing Knowledge Sources for Open-Domain Scientific Claim Verification
par: Vladika, Juraj, et autres
Publié: (2024)
par: Vladika, Juraj, et autres
Publié: (2024)
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
par: Fonseca, Marcio, et autres
Publié: (2024)
par: Fonseca, Marcio, et autres
Publié: (2024)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
par: Ramesh, Krithika, et autres
Publié: (2025)
par: Ramesh, Krithika, et autres
Publié: (2025)
FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning
par: Dou, Shaoyu, et autres
Publié: (2025)
par: Dou, Shaoyu, et autres
Publié: (2025)
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data
par: Afzal, Anum, et autres
Publié: (2024)
par: Afzal, Anum, et autres
Publié: (2024)
Adapting Large Language Models to Domains via Reading Comprehension
par: Cheng, Daixuan, et autres
Publié: (2023)
par: Cheng, Daixuan, et autres
Publié: (2023)
1-Diffractor: Efficient and Utility-Preserving Text Obfuscation Leveraging Word-Level Metric Differential Privacy
par: Meisenbacher, Stephen, et autres
Publié: (2024)
par: Meisenbacher, Stephen, et autres
Publié: (2024)
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
par: Croxford, Emma, et autres
Publié: (2024)
par: Croxford, Emma, et autres
Publié: (2024)
Knowledge Planning in Large Language Models for Domain-Aligned Counseling Summarization
par: Srivastava, Aseem, et autres
Publié: (2024)
par: Srivastava, Aseem, et autres
Publié: (2024)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
par: Aly, Walid Mohamed, et autres
Publié: (2025)
par: Aly, Walid Mohamed, et autres
Publié: (2025)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
par: Pokharel, Rhitabrat, et autres
Publié: (2025)
par: Pokharel, Rhitabrat, et autres
Publié: (2025)
SocialEval: Evaluating Social Intelligence of Large Language Models
par: Zhou, Jinfeng, et autres
Publié: (2025)
par: Zhou, Jinfeng, et autres
Publié: (2025)
AICoderEval: Improving AI Domain Code Generation of Large Language Models
par: Xia, Yinghui, et autres
Publié: (2024)
par: Xia, Yinghui, et autres
Publié: (2024)
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization
par: Balde, Gunjan, et autres
Publié: (2024)
par: Balde, Gunjan, et autres
Publié: (2024)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
par: Kammakomati, Mehant, et autres
Publié: (2024)
par: Kammakomati, Mehant, et autres
Publié: (2024)
Documents similaires
-
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
par: Mascarell, Laura, et autres
Publié: (2024) -
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
par: Afzal, Anum, et autres
Publié: (2025) -
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
par: Afzal, Anum, et autres
Publié: (2025) -
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
par: Afzal, Anum, et autres
Publié: (2025) -
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
par: Afzal, Anum, et autres
Publié: (2024)