Identifying Reliable Evaluation Metrics for Scientific Text Revision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jourdan, Léane, Boudin, Florian, Dufour, Richard, Hernandez, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text revision in Scientific Writing Assistance: An Overview
von: Jourdan, Léane, et al.
Veröffentlicht: (2023)
von: Jourdan, Léane, et al.
Veröffentlicht: (2023)
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
von: Jourdan, Leane, et al.
Veröffentlicht: (2024)
von: Jourdan, Leane, et al.
Veröffentlicht: (2024)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
von: Jourdan, Léane, et al.
Veröffentlicht: (2025)
von: Jourdan, Léane, et al.
Veröffentlicht: (2025)
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
von: Jourdan, Léane, et al.
Veröffentlicht: (2026)
von: Jourdan, Léane, et al.
Veröffentlicht: (2026)
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
ACL-rlg: A Dataset for Reading List Generation
von: Aubert-Béduchaud, Julien, et al.
Veröffentlicht: (2024)
von: Aubert-Béduchaud, Julien, et al.
Veröffentlicht: (2024)
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
Evaluating the Homogeneity of Keyphrase Prediction Models
von: Houbre, Maël, et al.
Veröffentlicht: (2026)
von: Houbre, Maël, et al.
Veröffentlicht: (2026)
Self-Compositional Data Augmentation for Scientific Keyphrase Generation
von: Houbre, Mael, et al.
Veröffentlicht: (2024)
von: Houbre, Mael, et al.
Veröffentlicht: (2024)
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
von: Boudin, Florian, et al.
Veröffentlicht: (2024)
von: Boudin, Florian, et al.
Veröffentlicht: (2024)
Preface to the Special Issue of the TAL Journal on Scholarly Document Processing
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
A Survey of Pre-trained Language Models for Processing Scientific Text
von: Ho, Xanh, et al.
Veröffentlicht: (2024)
von: Ho, Xanh, et al.
Veröffentlicht: (2024)
*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
von: Lemesle, Quentin, et al.
Veröffentlicht: (2026)
von: Lemesle, Quentin, et al.
Veröffentlicht: (2026)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
von: Kumar, Sunisth, et al.
Veröffentlicht: (2026)
von: Kumar, Sunisth, et al.
Veröffentlicht: (2026)
Automatically Suggesting Diverse Example Sentences for L2 Japanese Learners Using Pre-Trained Language Models
von: Benedetti, Enrico, et al.
Veröffentlicht: (2025)
von: Benedetti, Enrico, et al.
Veröffentlicht: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics
von: Roux, Thibault Bañeras, et al.
Veröffentlicht: (2026)
von: Roux, Thibault Bañeras, et al.
Veröffentlicht: (2026)
Renard: A Modular Pipeline for Extracting Character Networks from Narrative Texts
von: Amalvy, Arthur, et al.
Veröffentlicht: (2024)
von: Amalvy, Arthur, et al.
Veröffentlicht: (2024)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
von: Jourdan, Fanny
Veröffentlicht: (2024)
von: Jourdan, Fanny
Veröffentlicht: (2024)
Language Model Adaptation to Specialized Domains through Selective Masking based on Genre and Topical Characteristics
von: Belfathi, Anas, et al.
Veröffentlicht: (2024)
von: Belfathi, Anas, et al.
Veröffentlicht: (2024)
Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling
von: Belfathi, Anas, et al.
Veröffentlicht: (2026)
von: Belfathi, Anas, et al.
Veröffentlicht: (2026)
MoreHopQA: More Than Multi-hop Reasoning
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
Rethinking Scientific Summarization Evaluation: Grounding Explainable Metrics on Facet-aware Benchmark
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
Hacking Neural Evaluation Metrics with Single Hub Text
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2025)
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2025)
FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
von: Junqueras, Juan, et al.
Veröffentlicht: (2026)
von: Junqueras, Juan, et al.
Veröffentlicht: (2026)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
von: Ma, Yongqiang, et al.
Veröffentlicht: (2024)
von: Ma, Yongqiang, et al.
Veröffentlicht: (2024)
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
von: Afzal, Anum, et al.
Veröffentlicht: (2025)
von: Afzal, Anum, et al.
Veröffentlicht: (2025)
Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
von: Bañeras-Roux, Thibault, et al.
Veröffentlicht: (2026)
Reference-free Evaluation Metrics for Text Generation: A Survey
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
Evaluation Metrics for Text Data Augmentation in NLP
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
1-Diffractor: Efficient and Utility-Preserving Text Obfuscation Leveraging Word-Level Metric Differential Privacy
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
von: Jourdan, Fanny, et al.
Veröffentlicht: (2025)
von: Jourdan, Fanny, et al.
Veröffentlicht: (2025)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
von: Lorandi, Michela, et al.
Veröffentlicht: (2024)
von: Lorandi, Michela, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Text revision in Scientific Writing Assistance: An Overview
von: Jourdan, Léane, et al.
Veröffentlicht: (2023) -
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
von: Jourdan, Leane, et al.
Veröffentlicht: (2024) -
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
von: Jourdan, Léane, et al.
Veröffentlicht: (2025) -
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
von: Jourdan, Léane, et al.
Veröffentlicht: (2026) -
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
von: Boudin, Florian, et al.
Veröffentlicht: (2025)