The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks
Fuente:
arXiv
Saved in:
| Main Authors: | Hopton, Zachary, Vamvas, Jannis, Büchler, Andrin, Rutkiewicz, Anna, Cathomas, Rico, Sennrich, Rico |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RUMLEM: A Dictionary-Based Lemmatizer for Romansh
by: Fischer, Dominic P., et al.
Published: (2026)
by: Fischer, Dominic P., et al.
Published: (2026)
Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties
by: Vamvas, Jannis, et al.
Published: (2026)
by: Vamvas, Jannis, et al.
Published: (2026)
Robust Language Identification for Romansh Varieties
by: Model, Charlotte, et al.
Published: (2026)
by: Model, Charlotte, et al.
Published: (2026)
Linear-time Minimum Bayes Risk Decoding with Reference Aggregation
by: Vamvas, Jannis, et al.
Published: (2024)
by: Vamvas, Jannis, et al.
Published: (2024)
20min-XD: A Comparable Corpus of Swiss News Articles
by: Wastl, Michelle, et al.
Published: (2025)
by: Wastl, Michelle, et al.
Published: (2025)
SwissBERT: The Multilingual Language Model for Switzerland
by: Vamvas, Jannis, et al.
Published: (2023)
by: Vamvas, Jannis, et al.
Published: (2023)
Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents
by: Hu, Hanxu, et al.
Published: (2025)
by: Hu, Hanxu, et al.
Published: (2025)
Mitigating Hallucinations and Off-target Machine Translation with Source-Contrastive and Language-Contrastive Decoding
by: Sennrich, Rico, et al.
Published: (2023)
by: Sennrich, Rico, et al.
Published: (2023)
SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents
by: Wastl, Michelle, et al.
Published: (2025)
by: Wastl, Michelle, et al.
Published: (2025)
Machine Translation Models are Zero-Shot Detectors of Translation Direction
by: Wastl, Michelle, et al.
Published: (2024)
by: Wastl, Michelle, et al.
Published: (2024)
Modular Adaptation of Multilingual Encoders to Written Swiss German Dialect
by: Vamvas, Jannis, et al.
Published: (2024)
by: Vamvas, Jannis, et al.
Published: (2024)
Investigating Multi-Pivot Ensembling with Massively Multilingual Machine Translation Models
by: Mohammadshahi, Alireza, et al.
Published: (2023)
by: Mohammadshahi, Alireza, et al.
Published: (2023)
Leveraging In-Context Learning for Political Bias Testing of LLMs
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
QueST: Incentivizing LLMs to Generate Difficult Problems
by: Hu, Hanxu, et al.
Published: (2025)
by: Hu, Hanxu, et al.
Published: (2025)
DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning
by: Hu, Hanxu, et al.
Published: (2026)
by: Hu, Hanxu, et al.
Published: (2026)
Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
by: Semenov, Kirill, et al.
Published: (2025)
by: Semenov, Kirill, et al.
Published: (2025)
Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
by: Grosjean, Juri, et al.
Published: (2024)
by: Grosjean, Juri, et al.
Published: (2024)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
by: Michail, Andrianos, et al.
Published: (2025)
by: Michail, Andrianos, et al.
Published: (2025)
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
by: Kew, Tannon, et al.
Published: (2023)
by: Kew, Tannon, et al.
Published: (2023)
Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader
by: Vamvas, Jannis, et al.
Published: (2025)
by: Vamvas, Jannis, et al.
Published: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
by: Wu, Guojun, et al.
Published: (2024)
by: Wu, Guojun, et al.
Published: (2024)
Modeling Orthographic Variation in Occitan's Dialects
by: Hopton, Zachary William, et al.
Published: (2024)
by: Hopton, Zachary William, et al.
Published: (2024)
Robust Native Language Identification through Agentic Decomposition
by: Uluslu, Ahmet Yavuz, et al.
Published: (2025)
by: Uluslu, Ahmet Yavuz, et al.
Published: (2025)
Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
by: Schuhmacher, Elias, et al.
Published: (2026)
by: Schuhmacher, Elias, et al.
Published: (2026)
SignCLIP: Connecting Text and Sign Language by Contrastive Learning
by: Jiang, Zifan, et al.
Published: (2024)
by: Jiang, Zifan, et al.
Published: (2024)
Structure-Conditional Minimum Bayes Risk Decoding
by: Eikema, Bryan, et al.
Published: (2025)
by: Eikema, Bryan, et al.
Published: (2025)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language
by: Sennrich, Kilian, et al.
Published: (2025)
by: Sennrich, Kilian, et al.
Published: (2025)
CommonMorph: Participatory Morphological Documentation Platform
by: Mahmudi, Aso, et al.
Published: (2026)
by: Mahmudi, Aso, et al.
Published: (2026)
Phonetic stability across time: Linguistic enclaves in Switzerland*
by: Andrin Büchler
Published: (2022)
by: Andrin Büchler
Published: (2022)
MultimodalHugs: Enabling Sign Language Processing in Hugging Face
by: Sant, Gerard, et al.
Published: (2025)
by: Sant, Gerard, et al.
Published: (2025)
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models
by: Wang, Junling, et al.
Published: (2025)
by: Wang, Junling, et al.
Published: (2025)
Does mBERT understand Romansh? Evaluating word embeddings using word alignment
by: Dolev, Eyal Liron
Published: (2023)
by: Dolev, Eyal Liron
Published: (2023)
Parallel Corpus Augmentation using Masked Language Models
by: Kumari, Vibhuti, et al.
Published: (2024)
by: Kumari, Vibhuti, et al.
Published: (2024)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
Meaningful Pose-Based Sign Language Evaluation
by: Jiang, Zifan, et al.
Published: (2025)
by: Jiang, Zifan, et al.
Published: (2025)
Benchmarking POS Tagging for the Tajik Language: A Comparative Study of Neural Architectures on the TajPersParallel Corpus
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
EthioMT: Parallel Corpus for Low-resource Ethiopian Languages
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
GUMBridge: a Corpus for Varieties of Bridging Anaphora
by: Levine, Lauren, et al.
Published: (2025)
by: Levine, Lauren, et al.
Published: (2025)
Similar Items
-
RUMLEM: A Dictionary-Based Lemmatizer for Romansh
by: Fischer, Dominic P., et al.
Published: (2026) -
Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties
by: Vamvas, Jannis, et al.
Published: (2026) -
Robust Language Identification for Romansh Varieties
by: Model, Charlotte, et al.
Published: (2026) -
Linear-time Minimum Bayes Risk Decoding with Reference Aggregation
by: Vamvas, Jannis, et al.
Published: (2024) -
20min-XD: A Comparable Corpus of Swiss News Articles
by: Wastl, Michelle, et al.
Published: (2025)