Building and Aligning Comparable Corpora
Fuente:
arXiv
Saved in:
| Main Authors: | Saad, Motaz, Langlois, David, Smaili, Kamel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-lingual Opinions and Emotions Mining in Comparable Documents
by: Saad, Motaz, et al.
Published: (2025)
by: Saad, Motaz, et al.
Published: (2025)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025)
by: Doghmash, Salam Thabet, et al.
Published: (2025)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
by: Hahn, Udo
Published: (2024)
by: Hahn, Udo
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Aligning the Norwegian UD Treebank with Entity and Coreference Information
by: Jørgensen, Tollef Emil, et al.
Published: (2023)
by: Jørgensen, Tollef Emil, et al.
Published: (2023)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
by: Zhao, Yong, et al.
Published: (2025)
by: Zhao, Yong, et al.
Published: (2025)
IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages
by: VP, Sumesh
Published: (2026)
by: VP, Sumesh
Published: (2026)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
by: Lee, Lung-Hao, et al.
Published: (2026)
by: Lee, Lung-Hao, et al.
Published: (2026)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
Comparing Complex Concepts with Transformers: Matching Patent Claims Against Natural Language Text
by: Blume, Matthias, et al.
Published: (2024)
by: Blume, Matthias, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
by: Dang, Thao Anh, et al.
Published: (2024)
by: Dang, Thao Anh, et al.
Published: (2024)
UM_FHS at the CLEF 2025 SimpleText Track: Comparing No-Context and Fine-Tune Approaches for GPT-4.1 Models in Sentence and Document-Level Text Simplification
by: Kocbek, Primoz, et al.
Published: (2025)
by: Kocbek, Primoz, et al.
Published: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
by: Dhasmana, Akriti, et al.
Published: (2026)
by: Dhasmana, Akriti, et al.
Published: (2026)
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025)
by: Pohl, David, et al.
Published: (2025)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025)
by: Orlikowski, Matthias, et al.
Published: (2025)
Graphemic Normalization of the Perso-Arabic Script
by: Doctor, Raiomond, et al.
Published: (2022)
by: Doctor, Raiomond, et al.
Published: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
by: Peláez-González, Carlos, et al.
Published: (2025)
by: Peláez-González, Carlos, et al.
Published: (2025)
PaperVoyager : Building Interactive Web with Visual Language Models
by: Dai, Dasen, et al.
Published: (2026)
by: Dai, Dasen, et al.
Published: (2026)
Reliable Part-of-Speech Tagging of Historical Corpora through Set-Valued Prediction
by: Heid, Stefan, et al.
Published: (2020)
by: Heid, Stefan, et al.
Published: (2020)
PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
by: Sharma, Shivam, et al.
Published: (2025)
by: Sharma, Shivam, et al.
Published: (2025)
Evaluating Large Language Models for Zero-Shot Disease Labeling in CT Radiology Reports Across Organ Systems
by: Garcia-Alcoser, Michael E., et al.
Published: (2025)
by: Garcia-Alcoser, Michael E., et al.
Published: (2025)
Building Entity Association Mining Framework for Knowledge Discovery
by: Rawal, Anshika, et al.
Published: (2025)
by: Rawal, Anshika, et al.
Published: (2025)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
by: Nzeyimana, Antoine, et al.
Published: (2025)
by: Nzeyimana, Antoine, et al.
Published: (2025)
Low-resource neural machine translation with morphological modeling
by: Nzeyimana, Antoine
Published: (2024)
by: Nzeyimana, Antoine
Published: (2024)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
by: Li, Shanghao, et al.
Published: (2025)
by: Li, Shanghao, et al.
Published: (2025)
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
by: The Omnilingual MT Team, et al.
Published: (2025)
by: The Omnilingual MT Team, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Mitigating Translationese in Low-resource Languages: The Storyboard Approach
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
On Fusing ChatGPT and Ensemble Learning in Discon-tinuous Named Entity Recognition in Health Corpora
by: Chen, Tzu-Chieh, et al.
Published: (2024)
by: Chen, Tzu-Chieh, et al.
Published: (2024)
Omnilingual MT: Machine Translation for 1,600 Languages
by: Omnilingual MT Team, et al.
Published: (2026)
by: Omnilingual MT Team, et al.
Published: (2026)
A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
by: Park, Chiwan, et al.
Published: (2025)
by: Park, Chiwan, et al.
Published: (2025)
Similar Items
-
Cross-lingual Opinions and Emotions Mining in Comparable Documents
by: Saad, Motaz, et al.
Published: (2025) -
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025) -
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026) -
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
by: Hahn, Udo
Published: (2024) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)