ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | Jayawardena, Lasal, Yapa, Prasan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parameter Efficient Diverse Paraphrase Generation Using Sequence-Level Knowledge Distillation
by: Jayawardena, Lasal, et al.
Published: (2024)
by: Jayawardena, Lasal, et al.
Published: (2024)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
by: Jung, Jaehun, et al.
Published: (2023)
by: Jung, Jaehun, et al.
Published: (2023)
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026)
by: Agrawal, Pulin, et al.
Published: (2026)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
by: Albalak, Alon, et al.
Published: (2025)
by: Albalak, Alon, et al.
Published: (2025)
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
Analyzing Persuasive Strategies in Meme Texts: A Fusion of Language Models with Paraphrase Enrichment
by: Nayak, Kota Shamanth Ramanath, et al.
Published: (2024)
by: Nayak, Kota Shamanth Ramanath, et al.
Published: (2024)
Action Controlled Paraphrasing
by: Shi, Ning, et al.
Published: (2024)
by: Shi, Ning, et al.
Published: (2024)
Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors
by: Yadav, Vikas, et al.
Published: (2024)
by: Yadav, Vikas, et al.
Published: (2024)
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
by: Loula, João, et al.
Published: (2025)
by: Loula, João, et al.
Published: (2025)
ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection
by: Khairallah, Ali, et al.
Published: (2025)
by: Khairallah, Ali, et al.
Published: (2025)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
by: Hillebrand, Lars, et al.
Published: (2024)
by: Hillebrand, Lars, et al.
Published: (2024)
Syntactic Control of Language Models by Posterior Inference
by: Xefteri, Vicky, et al.
Published: (2025)
by: Xefteri, Vicky, et al.
Published: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
by: Kostić, Bogdan, et al.
Published: (2026)
by: Kostić, Bogdan, et al.
Published: (2026)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025)
by: Chan, Willy, et al.
Published: (2025)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
The Saturation Point of Backtranslation in High Quality Low Resource English Gujarati Machine Translation
by: Arif, Arwa
Published: (2025)
by: Arif, Arwa
Published: (2025)
Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
by: Nguyen, Huu, et al.
Published: (2025)
by: Nguyen, Huu, et al.
Published: (2025)
LOLgorithm: Integrating Semantic,Syntactic and Contextual Elements for Humor Classification
by: Khurana, Tanisha, et al.
Published: (2024)
by: Khurana, Tanisha, et al.
Published: (2024)
Vidur: A Large-Scale Simulation Framework For LLM Inference
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
by: Juzek, Tom S., et al.
Published: (2024)
by: Juzek, Tom S., et al.
Published: (2024)
BookSQL: A Large Scale Text-to-SQL Dataset for Accounting Domain
by: Kumar, Rahul, et al.
Published: (2024)
by: Kumar, Rahul, et al.
Published: (2024)
ContextGPT: Infusing LLMs Knowledge into Neuro-Symbolic Activity Recognition Models
by: Arrotta, Luca, et al.
Published: (2024)
by: Arrotta, Luca, et al.
Published: (2024)
Dual Encoder: Exploiting the Potential of Syntactic and Semantic for Aspect Sentiment Triplet Extraction
by: Zhao, Xiaowei, et al.
Published: (2024)
by: Zhao, Xiaowei, et al.
Published: (2024)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2024)
by: Köksal, Abdullatif, et al.
Published: (2024)
AMPS: ASR with Multimodal Paraphrase Supervision
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
by: Yang, Zhuonan, et al.
Published: (2026)
by: Yang, Zhuonan, et al.
Published: (2026)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
by: Xu, Ran, et al.
Published: (2023)
by: Xu, Ran, et al.
Published: (2023)
Scaling DPPs for RAG: Density Meets Diversity
by: Sun, Xun, et al.
Published: (2026)
by: Sun, Xun, et al.
Published: (2026)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023)
by: Kirk, Robert, et al.
Published: (2023)
Optimizing Diversity and Quality through Base-Aligned Model Collaboration
by: Wang, Yichen, et al.
Published: (2025)
by: Wang, Yichen, et al.
Published: (2025)
Understanding the Quality-Diversity Trade-off in Diffusion Language Models
by: Buzzard, Zak
Published: (2025)
by: Buzzard, Zak
Published: (2025)
Token-Level LLM Collaboration via FusionRoute
by: Xiong, Nuoya, et al.
Published: (2026)
by: Xiong, Nuoya, et al.
Published: (2026)
Curriculum Learning with Quality-Driven Data Selection
by: Wu, Biao, et al.
Published: (2024)
by: Wu, Biao, et al.
Published: (2024)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
Similar Items
-
Parameter Efficient Diverse Paraphrase Generation Using Sequence-Level Knowledge Distillation
by: Jayawardena, Lasal, et al.
Published: (2024) -
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
by: Jung, Jaehun, et al.
Published: (2023) -
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026) -
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026) -
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
by: Albalak, Alon, et al.
Published: (2025)