Molyé: A Corpus-based Approach to Language Contact in Colonial France
Fuente:
arXiv
Guardado en:
| Autores principales: | Dent, Rasul, Janès, Juliette, Clérice, Thibault, Suarez, Pedro Ortiz, Sagot, Benoît |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KréyoLID From Language Identification Towards Language Mining
por: Dent, Rasul, et al.
Publicado: (2025)
por: Dent, Rasul, et al.
Publicado: (2025)
How Should We Model the Probability of a Language?
por: Dent, Rasul, et al.
Publicado: (2026)
por: Dent, Rasul, et al.
Publicado: (2026)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
por: Karamolegkou, Antonia, et al.
Publicado: (2026)
por: Karamolegkou, Antonia, et al.
Publicado: (2026)
Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts
por: Clérice, Thibault
Publicado: (2023)
por: Clérice, Thibault
Publicado: (2023)
Diachronic Document Dataset for Semantic Layout Analysis
por: Clérice, Thibault, et al.
Publicado: (2024)
por: Clérice, Thibault, et al.
Publicado: (2024)
A French Version of the OLDI Seed Corpus
por: Marmonier, Malik, et al.
Publicado: (2025)
por: Marmonier, Malik, et al.
Publicado: (2025)
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
por: Futeral, Matthieu, et al.
Publicado: (2024)
por: Futeral, Matthieu, et al.
Publicado: (2024)
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
por: Zebaze, Armel, et al.
Publicado: (2025)
por: Zebaze, Armel, et al.
Publicado: (2025)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
por: Riabi, Arij, et al.
Publicado: (2021)
por: Riabi, Arij, et al.
Publicado: (2021)
You Actually Look Twice At it (YALTAi): using an object detection approach instead of region segmentation within the Kraken engine
por: Clérice, Thibault
Publicado: (2022)
por: Clérice, Thibault
Publicado: (2022)
From Text to Source: Results in Detecting Large Language Model-Generated Content
por: Antoun, Wissam, et al.
Publicado: (2023)
por: Antoun, Wissam, et al.
Publicado: (2023)
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
por: Clérice, Thibault, et al.
Publicado: (2026)
por: Clérice, Thibault, et al.
Publicado: (2026)
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
por: Godey, Nathan, et al.
Publicado: (2024)
por: Godey, Nathan, et al.
Publicado: (2024)
Tree of Problems: Improving structured problem solving with compositionality
por: Zebaze, Armel, et al.
Publicado: (2024)
por: Zebaze, Armel, et al.
Publicado: (2024)
In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
por: Zebaze, Armel, et al.
Publicado: (2024)
por: Zebaze, Armel, et al.
Publicado: (2024)
Making Sentence Embeddings Robust to User-Generated Content
por: Nishimwe, Lydia, et al.
Publicado: (2024)
por: Nishimwe, Lydia, et al.
Publicado: (2024)
Testing the Deliteralization Hypothesis in Human and Machine Translation
por: Marmonier, Malik, et al.
Publicado: (2026)
por: Marmonier, Malik, et al.
Publicado: (2026)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
por: Antoun, Wissam, et al.
Publicado: (2025)
por: Antoun, Wissam, et al.
Publicado: (2025)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
por: Zebaze, Armel, et al.
Publicado: (2025)
por: Zebaze, Armel, et al.
Publicado: (2025)
Explicit Learning and the LLM in Machine Translation
por: Marmonier, Malik, et al.
Publicado: (2025)
por: Marmonier, Malik, et al.
Publicado: (2025)
Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
por: Marmonier, Malik, et al.
Publicado: (2026)
por: Marmonier, Malik, et al.
Publicado: (2026)
TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
por: Zebaze, Armel, et al.
Publicado: (2025)
por: Zebaze, Armel, et al.
Publicado: (2025)
Language-Switching Triggers Take a Latent Detour Through Language Models
por: Kulumba, Francis, et al.
Publicado: (2026)
por: Kulumba, Francis, et al.
Publicado: (2026)
On the Scaling Laws of Geographical Representation in Language Models
por: Godey, Nathan, et al.
Publicado: (2024)
por: Godey, Nathan, et al.
Publicado: (2024)
Disentangling meaning from language in LLM-based machine translation
por: Lasnier, Théo, et al.
Publicado: (2026)
por: Lasnier, Théo, et al.
Publicado: (2026)
Anisotropy Is Inherent to Self-Attention in Transformers
por: Godey, Nathan, et al.
Publicado: (2024)
por: Godey, Nathan, et al.
Publicado: (2024)
Towards Zero-Shot Multimodal Machine Translation
por: Futeral, Matthieu, et al.
Publicado: (2024)
por: Futeral, Matthieu, et al.
Publicado: (2024)
Kreyòl-MT: Building MT for Latin American, Caribbean and Colonial African Creole Languages
por: Robinson, Nathaniel R., et al.
Publicado: (2024)
por: Robinson, Nathaniel R., et al.
Publicado: (2024)
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
por: Bafna, Niyati, et al.
Publicado: (2023)
por: Bafna, Niyati, et al.
Publicado: (2023)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
por: Antoun, Wissam, et al.
Publicado: (2024)
por: Antoun, Wissam, et al.
Publicado: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
por: Chambon, Pierre, et al.
Publicado: (2025)
por: Chambon, Pierre, et al.
Publicado: (2025)
Gaperon: A Peppered English-French Generative Language Model Suite
por: Godey, Nathan, et al.
Publicado: (2025)
por: Godey, Nathan, et al.
Publicado: (2025)
PatentEval: Understanding Errors in Patent Generation
por: Zuo, You, et al.
Publicado: (2024)
por: Zuo, You, et al.
Publicado: (2024)
Patent Representation Learning via Self-supervision
por: Zuo, You, et al.
Publicado: (2025)
por: Zuo, You, et al.
Publicado: (2025)
A Computational Approach to Language Contact -- A Case Study of Persian
por: Basirat, Ali, et al.
Publicado: (2026)
por: Basirat, Ali, et al.
Publicado: (2026)
The TUB Sign Language Corpus Collection
por: Avramidis, Eleftherios, et al.
Publicado: (2025)
por: Avramidis, Eleftherios, et al.
Publicado: (2025)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
por: Knill, Kate, et al.
Publicado: (2024)
por: Knill, Kate, et al.
Publicado: (2024)
A Benchmark Corpus and Neural Approach for Sanskrit Derivative Nouns Analysis
por: Singh, Arun Kumar, et al.
Publicado: (2020)
por: Singh, Arun Kumar, et al.
Publicado: (2020)
Corpus-Based Approaches to Igbo Diacritic Restoration
por: Ezeani, Ignatius
Publicado: (2026)
por: Ezeani, Ignatius
Publicado: (2026)
Probing Omissions and Distortions in Transformer-based RDF-to-Text Models
por: Faille, Juliette, et al.
Publicado: (2024)
por: Faille, Juliette, et al.
Publicado: (2024)
Ejemplares similares
-
KréyoLID From Language Identification Towards Language Mining
por: Dent, Rasul, et al.
Publicado: (2025) -
How Should We Model the Probability of a Language?
por: Dent, Rasul, et al.
Publicado: (2026) -
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
por: Karamolegkou, Antonia, et al.
Publicado: (2026) -
Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts
por: Clérice, Thibault
Publicado: (2023) -
Diachronic Document Dataset for Semantic Layout Analysis
por: Clérice, Thibault, et al.
Publicado: (2024)