ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
Fuente:
arXiv
Saved in:
| Main Authors: | Jourdan, Léane, Hernandez, Nicolas, Dufour, Richard, Boudin, Florian, Aizawa, Akiko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying Reliable Evaluation Metrics for Scientific Text Revision
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
Text revision in Scientific Writing Assistance: An Overview
by: Jourdan, Léane, et al.
Published: (2023)
by: Jourdan, Léane, et al.
Published: (2023)
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
by: Jourdan, Leane, et al.
Published: (2024)
by: Jourdan, Leane, et al.
Published: (2024)
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
by: Jourdan, Léane, et al.
Published: (2026)
by: Jourdan, Léane, et al.
Published: (2026)
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
by: Boudin, Florian, et al.
Published: (2024)
by: Boudin, Florian, et al.
Published: (2024)
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
by: Boudin, Florian, et al.
Published: (2025)
by: Boudin, Florian, et al.
Published: (2025)
Preface to the Special Issue of the TAL Journal on Scholarly Document Processing
by: Boudin, Florian, et al.
Published: (2025)
by: Boudin, Florian, et al.
Published: (2025)
Self-Compositional Data Augmentation for Scientific Keyphrase Generation
by: Houbre, Mael, et al.
Published: (2024)
by: Houbre, Mael, et al.
Published: (2024)
Automatically Suggesting Diverse Example Sentences for L2 Japanese Learners Using Pre-Trained Language Models
by: Benedetti, Enrico, et al.
Published: (2025)
by: Benedetti, Enrico, et al.
Published: (2025)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
by: Ho, Xanh, et al.
Published: (2025)
by: Ho, Xanh, et al.
Published: (2025)
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
by: Kumar, Sunisth, et al.
Published: (2026)
by: Kumar, Sunisth, et al.
Published: (2026)
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
by: Ho, Xanh, et al.
Published: (2025)
by: Ho, Xanh, et al.
Published: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
by: Ho, Xanh, et al.
Published: (2026)
by: Ho, Xanh, et al.
Published: (2026)
MoreHopQA: More Than Multi-hop Reasoning
by: Schnitzler, Julian, et al.
Published: (2024)
by: Schnitzler, Julian, et al.
Published: (2024)
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
by: Ho, Xanh, et al.
Published: (2025)
by: Ho, Xanh, et al.
Published: (2025)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
by: Kim, Kon Woo, et al.
Published: (2025)
by: Kim, Kon Woo, et al.
Published: (2025)
A Survey of Pre-trained Language Models for Processing Scientific Text
by: Ho, Xanh, et al.
Published: (2024)
by: Ho, Xanh, et al.
Published: (2024)
ACL-rlg: A Dataset for Reading List Generation
by: Aubert-Béduchaud, Julien, et al.
Published: (2024)
by: Aubert-Béduchaud, Julien, et al.
Published: (2024)
Are Emotions Arranged in a Circle? Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning
by: Yamauchi, Yusuke, et al.
Published: (2026)
by: Yamauchi, Yusuke, et al.
Published: (2026)
FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
by: Junqueras, Juan, et al.
Published: (2026)
by: Junqueras, Juan, et al.
Published: (2026)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification
by: Li, Xiangci, et al.
Published: (2020)
by: Li, Xiangci, et al.
Published: (2020)
Evaluating the Homogeneity of Keyphrase Prediction Models
by: Houbre, Maël, et al.
Published: (2026)
by: Houbre, Maël, et al.
Published: (2026)
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
by: Jiang, Junfeng, et al.
Published: (2024)
by: Jiang, Junfeng, et al.
Published: (2024)
A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific Publications
by: Jeangirard, Eric
Published: (2025)
by: Jeangirard, Eric
Published: (2025)
*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
by: Lemesle, Quentin, et al.
Published: (2026)
by: Lemesle, Quentin, et al.
Published: (2026)
Tracing Multilingual Knowledge Acquisition Dynamics in Domain Adaptation: A Case Study of English-Japanese Biomedical Adaptation
by: Zhao, Xin, et al.
Published: (2025)
by: Zhao, Xin, et al.
Published: (2025)
Refining and Reusing Annotation Guidelines for LLM Annotation
by: Kim, Kon Woo, et al.
Published: (2026)
by: Kim, Kon Woo, et al.
Published: (2026)
SKT5SciSumm -- Revisiting Extractive-Generative Approach for Multi-Document Scientific Summarization
by: To, Huy Quoc, et al.
Published: (2024)
by: To, Huy Quoc, et al.
Published: (2024)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
Extracting Paragraphs from LLM Token Activations
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature
by: Dao, An, et al.
Published: (2026)
by: Dao, An, et al.
Published: (2026)
Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering
by: Zhu, Yihua, et al.
Published: (2025)
by: Zhu, Yihua, et al.
Published: (2025)
Med-CoDE: Medical Critique based Disagreement Evaluation Framework
by: Gupta, Mohit, et al.
Published: (2025)
by: Gupta, Mohit, et al.
Published: (2025)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
by: Jourdan, Fanny
Published: (2024)
by: Jourdan, Fanny
Published: (2024)
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Through the LLM Looking Glass: A Socratic Probing of Donkeys, Elephants, and Markets
by: Kennedy, Molly, et al.
Published: (2025)
by: Kennedy, Molly, et al.
Published: (2025)
Language Model Adaptation to Specialized Domains through Selective Masking based on Genre and Topical Characteristics
by: Belfathi, Anas, et al.
Published: (2024)
by: Belfathi, Anas, et al.
Published: (2024)
Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling
by: Belfathi, Anas, et al.
Published: (2026)
by: Belfathi, Anas, et al.
Published: (2026)
Birbal: An efficient 7B instruct-model fine-tuned with curated datasets
by: Jindal, Ashvini Kumar, et al.
Published: (2024)
by: Jindal, Ashvini Kumar, et al.
Published: (2024)
Similar Items
-
Identifying Reliable Evaluation Metrics for Scientific Text Revision
by: Jourdan, Léane, et al.
Published: (2025) -
Text revision in Scientific Writing Assistance: An Overview
by: Jourdan, Léane, et al.
Published: (2023) -
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
by: Jourdan, Leane, et al.
Published: (2024) -
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
by: Jourdan, Léane, et al.
Published: (2026) -
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
by: Boudin, Florian, et al.
Published: (2024)