Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lorandi, Michela, Belz, Anya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comparative Study of Controlled Text Generation Systems Using Level-Playing-Field Evaluation Principles
von: Lorandi, Michela, et al.
Veröffentlicht: (2026)
von: Lorandi, Michela, et al.
Veröffentlicht: (2026)
Output Composability of QLoRA PEFT Modules for Plug-and-Play Attribute-Controlled Text Generation
von: Lorandi, Michela, et al.
Veröffentlicht: (2026)
von: Lorandi, Michela, et al.
Veröffentlicht: (2026)
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models
von: Lorandi, Michela, et al.
Veröffentlicht: (2024)
von: Lorandi, Michela, et al.
Veröffentlicht: (2024)
QRA++: Quantified Reproducibility Assessment for Common Types of Results in Natural Language Processing
von: Belz, Anya
Veröffentlicht: (2025)
von: Belz, Anya
Veröffentlicht: (2025)
Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)
HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
von: Belz, Anya, et al.
Veröffentlicht: (2024)
von: Belz, Anya, et al.
Veröffentlicht: (2024)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
von: Sabry, Mohammed, et al.
Veröffentlicht: (2024)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2024)
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
von: Sabry, Mohammed, et al.
Veröffentlicht: (2025)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2025)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
von: Sabry, Mohammed, et al.
Veröffentlicht: (2026)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2026)
The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
von: Belz, Anya, et al.
Veröffentlicht: (2025)
von: Belz, Anya, et al.
Veröffentlicht: (2025)
News is More than a Collection of Facts: Moral Frame Preserving News Summarization
von: Liscio, Enrico, et al.
Veröffentlicht: (2025)
von: Liscio, Enrico, et al.
Veröffentlicht: (2025)
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2026)
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2026)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
Reference-free Evaluation Metrics for Text Generation: A Survey
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
Generic Embedding-Based Lexicons for Transparent and Reproducible Text Scoring
von: Moez, Catherine
Veröffentlicht: (2024)
von: Moez, Catherine
Veröffentlicht: (2024)
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
von: Yang, Yi, et al.
Veröffentlicht: (2024)
von: Yang, Yi, et al.
Veröffentlicht: (2024)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
von: Dejl, Adam, et al.
Veröffentlicht: (2025)
von: Dejl, Adam, et al.
Veröffentlicht: (2025)
Identifying Reliable Evaluation Metrics for Scientific Text Revision
von: Jourdan, Léane, et al.
Veröffentlicht: (2025)
von: Jourdan, Léane, et al.
Veröffentlicht: (2025)
Hacking Neural Evaluation Metrics with Single Hub Text
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2025)
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2025)
Evaluation Metrics for Text Data Augmentation in NLP
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)
TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation
von: Grimal, Paul, et al.
Veröffentlicht: (2023)
von: Grimal, Paul, et al.
Veröffentlicht: (2023)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
Reproducing Complex Set-Compositional Information Retrieval
von: Degenhart, Vincent, et al.
Veröffentlicht: (2026)
von: Degenhart, Vincent, et al.
Veröffentlicht: (2026)
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
von: Zhou, Shang, et al.
Veröffentlicht: (2024)
von: Zhou, Shang, et al.
Veröffentlicht: (2024)
Evaluating Prompt Engineering Strategies for Sentiment Control in AI-Generated Texts
von: Sahler, Kerstin, et al.
Veröffentlicht: (2026)
von: Sahler, Kerstin, et al.
Veröffentlicht: (2026)
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
Faithful Model Evaluation for Model-Based Metrics
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
Calibrating Model-Based Evaluation Metrics for Summarization
von: Liu, Hongye, et al.
Veröffentlicht: (2026)
von: Liu, Hongye, et al.
Veröffentlicht: (2026)
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
von: Chen, Xiuying, et al.
Veröffentlicht: (2025)
von: Chen, Xiuying, et al.
Veröffentlicht: (2025)
GAMBIT+: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics
von: Filandrianos, Giorgos, et al.
Veröffentlicht: (2025)
von: Filandrianos, Giorgos, et al.
Veröffentlicht: (2025)
Scalable Parameter-Light Spectral Method for Clustering Short Text Embeddings with a Cohesion-Based Evaluation Metric
von: Neveditsin, Nikita, et al.
Veröffentlicht: (2025)
von: Neveditsin, Nikita, et al.
Veröffentlicht: (2025)
REFeREE: A REference-FREE Model-Based Metric for Text Simplification
von: Huang, Yichen, et al.
Veröffentlicht: (2024)
von: Huang, Yichen, et al.
Veröffentlicht: (2024)
Lessons from the Trenches on Reproducible Evaluation of Language Models
von: Biderman, Stella, et al.
Veröffentlicht: (2024)
von: Biderman, Stella, et al.
Veröffentlicht: (2024)
Nomic Embed: Training a Reproducible Long Context Text Embedder
von: Nussbaum, Zach, et al.
Veröffentlicht: (2024)
von: Nussbaum, Zach, et al.
Veröffentlicht: (2024)
A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text
von: Mikros, George, et al.
Veröffentlicht: (2026)
von: Mikros, George, et al.
Veröffentlicht: (2026)
NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization
von: Kafi, Md Abdullah Al, et al.
Veröffentlicht: (2025)
von: Kafi, Md Abdullah Al, et al.
Veröffentlicht: (2025)
Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
von: Jiang, Changhao, et al.
Veröffentlicht: (2026)
von: Jiang, Changhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Comparative Study of Controlled Text Generation Systems Using Level-Playing-Field Evaluation Principles
von: Lorandi, Michela, et al.
Veröffentlicht: (2026) -
Output Composability of QLoRA PEFT Modules for Plug-and-Play Attribute-Controlled Text Generation
von: Lorandi, Michela, et al.
Veröffentlicht: (2026) -
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models
von: Lorandi, Michela, et al.
Veröffentlicht: (2024) -
QRA++: Quantified Reproducibility Assessment for Common Types of Results in Natural Language Processing
von: Belz, Anya
Veröffentlicht: (2025) -
Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)