Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
Fuente:
arXiv
Saved in:
| Main Authors: | Mukherjee, Sourabrata, Ojha, Atul Kr., McCrae, John P., Dusek, Ondrej |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text Detoxification as Style Transfer in English and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Multilingual Text Style Transfer: Datasets & Models for Indian Languages
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Are Large Language Models Actually Good at Text Style Transfer?
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Text Style Transfer: An Introductory Overview
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
A Survey of Text Style Transfer: Applications and Ethical Implications
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Inferring Adjective Hypernyms with Language Models to Increase the Connectivity of Open English Wordnet
by: Augello, Lorenzo, et al.
Published: (2025)
by: Augello, Lorenzo, et al.
Published: (2025)
\textit{Versteasch du mi?} Computational and Socio-Linguistic Perspectives on GenAI, LLMs, and Non-Standard Language
by: Platzgummer, Verena, et al.
Published: (2026)
by: Platzgummer, Verena, et al.
Published: (2026)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)
by: Kartáč, Ivan, et al.
Published: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
by: Suri, Omar Farooq Khan, et al.
Published: (2025)
by: Suri, Omar Farooq Khan, et al.
Published: (2025)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
by: Onderková, Kristýna, et al.
Published: (2025)
by: Onderková, Kristýna, et al.
Published: (2025)
When retrieval outperforms generation: Dense evidence retrieval for scalable fake news detection
by: Qazi, Alamgir Munir, et al.
Published: (2025)
by: Qazi, Alamgir Munir, et al.
Published: (2025)
MaCmS: Magahi Code-mixed Dataset for Sentiment Analysis
by: Rani, Priya, et al.
Published: (2024)
by: Rani, Priya, et al.
Published: (2024)
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis
by: Negi, Gaurav, et al.
Published: (2026)
by: Negi, Gaurav, et al.
Published: (2026)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
Real-World Summarization: When Evaluation Reaches Its Limits
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
by: Negi, Gaurav, et al.
Published: (2025)
by: Negi, Gaurav, et al.
Published: (2025)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
by: Lango, Mateusz, et al.
Published: (2025)
by: Lango, Mateusz, et al.
Published: (2025)
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
by: Schmidtová, Patrícia, et al.
Published: (2024)
by: Schmidtová, Patrícia, et al.
Published: (2024)
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics
by: Pauli, Amalie Brogaard, et al.
Published: (2025)
by: Pauli, Amalie Brogaard, et al.
Published: (2025)
Comparing LLM prompting with Cross-lingual transfer performance on Indigenous and Low-resource Brazilian Languages
by: Adelani, David Ifeoluwa, et al.
Published: (2024)
by: Adelani, David Ifeoluwa, et al.
Published: (2024)
AnimatedLLM: Explaining LLMs with Interactive Visualizations
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems
by: Kumar, Nalin, et al.
Published: (2023)
by: Kumar, Nalin, et al.
Published: (2023)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
by: Warczyński, Jędrzej, et al.
Published: (2025)
by: Warczyński, Jędrzej, et al.
Published: (2025)
Identifying Reliable Evaluation Metrics for Scientific Text Revision
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
LMStyle Benchmark: Evaluating Text Style Transfer for Chatbots
by: Chen, Jianlin
Published: (2024)
by: Chen, Jianlin
Published: (2024)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
by: Balloccu, Simone, et al.
Published: (2024)
by: Balloccu, Simone, et al.
Published: (2024)
Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
by: Protasov, Vitaly, et al.
Published: (2025)
by: Protasov, Vitaly, et al.
Published: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
by: Kamzela, Wiktor, et al.
Published: (2025)
by: Kamzela, Wiktor, et al.
Published: (2025)
Strategies for Span Labeling with Large Language Models
by: Semin, Danil, et al.
Published: (2026)
by: Semin, Danil, et al.
Published: (2026)
Reasoning Gets Harder for LLMs Inside A Dialogue
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
Findings of the IWSLT 2024 Evaluation Campaign
by: Ahmad, Ibrahim Said, et al.
Published: (2024)
by: Ahmad, Ibrahim Said, et al.
Published: (2024)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
by: Javorský, Dávid, et al.
Published: (2022)
by: Javorský, Dávid, et al.
Published: (2022)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
by: Hamna, Hamna, et al.
Published: (2025)
by: Hamna, Hamna, et al.
Published: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024)
by: Wojciechowski, Adam, et al.
Published: (2024)
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Evaluating Style-Personalized Text Generation: Challenges and Directions
by: Jangra, Anubhav, et al.
Published: (2025)
by: Jangra, Anubhav, et al.
Published: (2025)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Similar Items
-
Text Detoxification as Style Transfer in English and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2024) -
Multilingual Text Style Transfer: Datasets & Models for Indian Languages
by: Mukherjee, Sourabrata, et al.
Published: (2024) -
Are Large Language Models Actually Good at Text Style Transfer?
by: Mukherjee, Sourabrata, et al.
Published: (2024) -
Text Style Transfer: An Introductory Overview
by: Mukherjee, Sourabrata, et al.
Published: (2024) -
A Survey of Text Style Transfer: Applications and Ethical Implications
by: Mukherjee, Sourabrata, et al.
Published: (2024)