SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Ziyao, Wang, Houfeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion
by: Xu, Ziyao, et al.
Published: (2025)
by: Xu, Ziyao, et al.
Published: (2025)
Detection-Correction Structure via General Language Model for Grammatical Error Correction
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective
by: Xu, Ziyao, et al.
Published: (2026)
by: Xu, Ziyao, et al.
Published: (2026)
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation
by: Luo, Wen, et al.
Published: (2025)
by: Luo, Wen, et al.
Published: (2025)
An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage
by: Bu, Fan, et al.
Published: (2025)
by: Bu, Fan, et al.
Published: (2025)
DEE: Dual-stage Explainable Evaluation Method for Text Generation
by: Zhang, Shenyu, et al.
Published: (2024)
by: Zhang, Shenyu, et al.
Published: (2024)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
by: Tao, Zhen, et al.
Published: (2024)
by: Tao, Zhen, et al.
Published: (2024)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
CiteCheck: Towards Accurate Citation Faithfulness Detection
by: Xu, Ziyao, et al.
Published: (2025)
by: Xu, Ziyao, et al.
Published: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation
by: Zhong, Tianqi, et al.
Published: (2024)
by: Zhong, Tianqi, et al.
Published: (2024)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
by: Ramesh, Krithika, et al.
Published: (2025)
by: Ramesh, Krithika, et al.
Published: (2025)
Structsum Generation for Faster Text Comprehension
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
by: Onderková, Kristýna, et al.
Published: (2025)
by: Onderková, Kristýna, et al.
Published: (2025)
A Comprehensive Dataset for Human vs. AI Generated Text Detection
by: Roy, Rajarshi, et al.
Published: (2025)
by: Roy, Rajarshi, et al.
Published: (2025)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities
by: Chai, Yaping, et al.
Published: (2025)
by: Chai, Yaping, et al.
Published: (2025)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
by: Luo, Wen, et al.
Published: (2024)
by: Luo, Wen, et al.
Published: (2024)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
by: Schoenegger, Loris, et al.
Published: (2024)
by: Schoenegger, Loris, et al.
Published: (2024)
An Extensive Evaluation of Factual Consistency in Large Language Models for Data-to-Text Generation
by: Mahapatra, Joy, et al.
Published: (2024)
by: Mahapatra, Joy, et al.
Published: (2024)
Advancements in Scientific Controllable Text Generation Methods
by: Goel, Arnav, et al.
Published: (2023)
by: Goel, Arnav, et al.
Published: (2023)
A Practical Method for Generating String Counterfactuals
by: Avitan, Matan, et al.
Published: (2024)
by: Avitan, Matan, et al.
Published: (2024)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
by: Ngo, Nghia Trung, et al.
Published: (2024)
by: Ngo, Nghia Trung, et al.
Published: (2024)
A Frustratingly Simple Decoding Method for Neural Text Generation
by: Yang, Haoran, et al.
Published: (2023)
by: Yang, Haoran, et al.
Published: (2023)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
by: Liu, Xiaoze, et al.
Published: (2024)
by: Liu, Xiaoze, et al.
Published: (2024)
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
by: Zhou, Shang, et al.
Published: (2024)
by: Zhou, Shang, et al.
Published: (2024)
Multilingual Hate Speech Detection and Counterspeech Generation: A Comprehensive Survey and Practical Guide
by: Fesaghandis, Zahra Safdari, et al.
Published: (2026)
by: Fesaghandis, Zahra Safdari, et al.
Published: (2026)
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
by: Chen, Xiuying, et al.
Published: (2025)
by: Chen, Xiuying, et al.
Published: (2025)
Data Swarms: Optimizable Generation of Synthetic Evaluation Data
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
Learning Personalized Alignment for Evaluating Open-ended Text Generation
by: Wang, Danqing, et al.
Published: (2023)
by: Wang, Danqing, et al.
Published: (2023)
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
Reference-free Evaluation Metrics for Text Generation: A Survey
by: Ito, Takumi, et al.
Published: (2025)
by: Ito, Takumi, et al.
Published: (2025)
LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods
by: Wang, Zhenhua, et al.
Published: (2024)
by: Wang, Zhenhua, et al.
Published: (2024)
Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
Similar Items
-
Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion
by: Xu, Ziyao, et al.
Published: (2025) -
Detection-Correction Structure via General Language Model for Grammatical Error Correction
by: Li, Wei, et al.
Published: (2024) -
Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
by: Wang, Zihan, et al.
Published: (2024) -
Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective
by: Xu, Ziyao, et al.
Published: (2026) -
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation
by: Luo, Wen, et al.
Published: (2025)