MASRAD: Arabic Terminology Management Corpora with Semi-Automatic Construction
Fuente:
arXiv
Saved in:
| Main Authors: | Nasser, Mahdi, Sayyah, Laura, Zaraket, Fadi A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Back-of-the-Book Index Automation for Arabic Documents
by: Haidar, Nawal, et al.
Published: (2024)
by: Haidar, Nawal, et al.
Published: (2024)
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
by: Hennara, Khalil, et al.
Published: (2025)
by: Hennara, Khalil, et al.
Published: (2025)
A Novel Dialect-Aware Framework for the Classification of Arabic Dialects and Emotions
by: Alsadhan, Nasser A
Published: (2025)
by: Alsadhan, Nasser A
Published: (2025)
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
by: Abbas, Chaymaa, et al.
Published: (2026)
by: Abbas, Chaymaa, et al.
Published: (2026)
AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts
by: Huang, Baorong, et al.
Published: (2025)
by: Huang, Baorong, et al.
Published: (2025)
Is AI Catching Up to Human Expression? Exploring Emotion, Personality, Authorship, and Linguistic Style in English and Arabic with Six Large Language Models
by: Alsadhan, Nasser A
Published: (2026)
by: Alsadhan, Nasser A
Published: (2026)
Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
How Well Do LLMs Understand Tunisian Arabic?
by: Mahdi, Mohamed
Published: (2025)
by: Mahdi, Mohamed
Published: (2025)
The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora
by: Amiraz, Chen, et al.
Published: (2025)
by: Amiraz, Chen, et al.
Published: (2025)
Domain Terminology Integration into Machine Translation: Leveraging Large Language Models
by: Moslem, Yasmin, et al.
Published: (2023)
by: Moslem, Yasmin, et al.
Published: (2023)
Is This Collection Worth My LLM's Time? Automatically Measuring Information Potential in Text Corpora
by: Karch, Tristan, et al.
Published: (2025)
by: Karch, Tristan, et al.
Published: (2025)
Aligning ESG Controversy Data with International Guidelines through Semi-Automatic Ontology Construction
by: Iwata, Tsuyoshi, et al.
Published: (2025)
by: Iwata, Tsuyoshi, et al.
Published: (2025)
Automatic Classification of Arabic Literature into Historical Eras
by: Alhathloul, Zainab, et al.
Published: (2026)
by: Alhathloul, Zainab, et al.
Published: (2026)
AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora
by: Wu, Xiaoqing, et al.
Published: (2026)
by: Wu, Xiaoqing, et al.
Published: (2026)
EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech
by: Hassan, Besher, et al.
Published: (2026)
by: Hassan, Besher, et al.
Published: (2026)
Automatic Construction of Multiple Classification Dimensions for Managing Approaches in Scientific Papers
by: Ma, Bing, et al.
Published: (2025)
by: Ma, Bing, et al.
Published: (2025)
Arabic Automatic Story Generation with Large Language Models
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
by: Uro, Rémi, et al.
Published: (2024)
by: Uro, Rémi, et al.
Published: (2024)
A Method for Learning Large-Scale Computational Construction Grammars from Semantically Annotated Corpora
by: Van Eecke, Paul, et al.
Published: (2026)
by: Van Eecke, Paul, et al.
Published: (2026)
CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic
by: Mankarious, Saad, et al.
Published: (2025)
by: Mankarious, Saad, et al.
Published: (2025)
TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora
by: Kargupta, Priyanka, et al.
Published: (2025)
by: Kargupta, Priyanka, et al.
Published: (2025)
Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
by: Gomes, Juliana Resplande Sant'anna, et al.
Published: (2025)
by: Gomes, Juliana Resplande Sant'anna, et al.
Published: (2025)
Validating and Exploring Large Geographic Corpora
by: Dunn, Jonathan
Published: (2024)
by: Dunn, Jonathan
Published: (2024)
Identifying Emerging Concepts in Large Corpora
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain
by: Obaidah, Qusai Abo, et al.
Published: (2024)
by: Obaidah, Qusai Abo, et al.
Published: (2024)
A Hierarchical and Attentional Analysis of Argument Structure Constructions in BERT Using Naturalistic Corpora
by: Kaipeng, Liu, et al.
Published: (2026)
by: Kaipeng, Liu, et al.
Published: (2026)
Toward Human-Centered AI-Assisted Terminology Work
by: Martin, Antonio San
Published: (2025)
by: Martin, Antonio San
Published: (2025)
KpopMT: Translation Dataset with Terminology for Kpop Fandom
by: Kim, JiWoo, et al.
Published: (2024)
by: Kim, JiWoo, et al.
Published: (2024)
Readability Measures and Automatic Text Simplification: In the Search of a Construct
by: Cardon, Rémi, et al.
Published: (2025)
by: Cardon, Rémi, et al.
Published: (2025)
From RAG to Agentic RAG for Faithful Islamic Question Answering
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
Comparable Corpora: Opportunities for New Research Directions
by: Church, Kenneth
Published: (2025)
by: Church, Kenneth
Published: (2025)
New Textual Corpora for Serbian Language Modeling
by: Škorić, Mihailo, et al.
Published: (2024)
by: Škorić, Mihailo, et al.
Published: (2024)
Bias in News Summarization: Measures, Pitfalls and Corpora
by: Steen, Julius, et al.
Published: (2023)
by: Steen, Julius, et al.
Published: (2023)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
Mathematical Entities: Corpora and Benchmarks
by: Collard, Jacob, et al.
Published: (2024)
by: Collard, Jacob, et al.
Published: (2024)
Automatic Construction of Chinese Verb Collostruction Database
by: Tang, Xuri, et al.
Published: (2025)
by: Tang, Xuri, et al.
Published: (2025)
Automatic Knowledge Graph Construction for Judicial Cases
by: Zhou, Jie, et al.
Published: (2024)
by: Zhou, Jie, et al.
Published: (2024)
Learning to Translate Ambiguous Terminology by Preference Optimization on Post-Edits
by: Berger, Nathaniel, et al.
Published: (2025)
by: Berger, Nathaniel, et al.
Published: (2025)
Lingua Custodi's participation at the WMT 2025 Terminology shared task
by: Liu, Jingshu, et al.
Published: (2025)
by: Liu, Jingshu, et al.
Published: (2025)
Efficient Terminology Integration for LLM-based Translation in Specialized Domains
by: Kim, Sejoon, et al.
Published: (2024)
by: Kim, Sejoon, et al.
Published: (2024)
Similar Items
-
Back-of-the-Book Index Automation for Arabic Documents
by: Haidar, Nawal, et al.
Published: (2024) -
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
by: Hennara, Khalil, et al.
Published: (2025) -
A Novel Dialect-Aware Framework for the Classification of Arabic Dialects and Emotions
by: Alsadhan, Nasser A
Published: (2025) -
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
by: Abbas, Chaymaa, et al.
Published: (2026) -
AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts
by: Huang, Baorong, et al.
Published: (2025)