Gaperon: A Peppered English-French Generative Language Model Suite
Fuente:
arXiv
Saved in:
| Main Authors: | Godey, Nathan, Antoun, Wissam, Touchent, Rian, Bawden, Rachel, de la Clergerie, Éric, Sagot, Benoît, Seddah, Djamé |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024)
by: Antoun, Wissam, et al.
Published: (2024)
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)
by: Antoun, Wissam, et al.
Published: (2023)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
by: Antoun, Wissam, et al.
Published: (2025)
by: Antoun, Wissam, et al.
Published: (2025)
On the Scaling Laws of Geographical Representation in Language Models
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
by: Touchent, Rian, et al.
Published: (2025)
by: Touchent, Rian, et al.
Published: (2025)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
by: Touchent, Rian, et al.
Published: (2026)
by: Touchent, Rian, et al.
Published: (2026)
Language-Switching Triggers Take a Latent Detour Through Language Models
by: Kulumba, Francis, et al.
Published: (2026)
by: Kulumba, Francis, et al.
Published: (2026)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
by: Touchent, Rian, et al.
Published: (2023)
by: Touchent, Rian, et al.
Published: (2023)
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Disentangling meaning from language in LLM-based machine translation
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021)
by: Riabi, Arij, et al.
Published: (2021)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
A French Version of the OLDI Seed Corpus
by: Marmonier, Malik, et al.
Published: (2025)
by: Marmonier, Malik, et al.
Published: (2025)
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
Making Sentence Embeddings Robust to User-Generated Content
by: Nishimwe, Lydia, et al.
Published: (2024)
by: Nishimwe, Lydia, et al.
Published: (2024)
PatentEval: Understanding Errors in Patent Generation
by: Zuo, You, et al.
Published: (2024)
by: Zuo, You, et al.
Published: (2024)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
Tree of Problems: Improving structured problem solving with compositionality
by: Zebaze, Armel, et al.
Published: (2024)
by: Zebaze, Armel, et al.
Published: (2024)
Testing the Deliteralization Hypothesis in Human and Machine Translation
by: Marmonier, Malik, et al.
Published: (2026)
by: Marmonier, Malik, et al.
Published: (2026)
In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2024)
by: Zebaze, Armel, et al.
Published: (2024)
Explicit Learning and the LLM in Machine Translation
by: Marmonier, Malik, et al.
Published: (2025)
by: Marmonier, Malik, et al.
Published: (2025)
Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
by: Marmonier, Malik, et al.
Published: (2026)
by: Marmonier, Malik, et al.
Published: (2026)
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
by: Mouilleron, Virginie, et al.
Published: (2026)
by: Mouilleron, Virginie, et al.
Published: (2026)
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
by: Riabi, Arij, et al.
Published: (2023)
by: Riabi, Arij, et al.
Published: (2023)
Towards Zero-Shot Multimodal Machine Translation
by: Futeral, Matthieu, et al.
Published: (2024)
by: Futeral, Matthieu, et al.
Published: (2024)
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
by: Lopetegui, Javier A., et al.
Published: (2024)
by: Lopetegui, Javier A., et al.
Published: (2024)
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
by: Bafna, Niyati, et al.
Published: (2023)
by: Bafna, Niyati, et al.
Published: (2023)
Patent Representation Learning via Self-supervision
by: Zuo, You, et al.
Published: (2025)
by: Zuo, You, et al.
Published: (2025)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025)
by: Chambon, Pierre, et al.
Published: (2025)
Lost in Backpropagation: The LM Head is a Gradient Bottleneck
by: Godey, Nathan, et al.
Published: (2026)
by: Godey, Nathan, et al.
Published: (2026)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)
by: Karamolegkou, Antonia, et al.
Published: (2026)
Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America
by: Karmim, Yannis, et al.
Published: (2026)
by: Karmim, Yannis, et al.
Published: (2026)
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
by: Clérice, Thibault, et al.
Published: (2026)
by: Clérice, Thibault, et al.
Published: (2026)
Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
Similar Items
-
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024) -
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023) -
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
by: Antoun, Wissam, et al.
Published: (2025) -
On the Scaling Laws of Geographical Representation in Language Models
by: Godey, Nathan, et al.
Published: (2024) -
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
by: Touchent, Rian, et al.
Published: (2025)