:: Library Catalog

Cover Image

Saved in:

Bibliographic Details
Main Authors:	Nishimwe, Lydia, Sagot, Benoît, Bawden, Rachel
Format:	Preprint
Published:	2024
Subjects:	Computation and Language
Online Access:	https://arxiv.org/abs/2403.17220
Tags:	Add Tag No Tags, Be the first to tag this record!

Similar Items

LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
by: Zebaze, Armel, et al.
Published: (2025)

TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)

Tree of Problems: Improving structured problem solving with compositionality
by: Zebaze, Armel, et al.
Published: (2024)

In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2024)

A French Version of the OLDI Seed Corpus
by: Marmonier, Malik, et al.
Published: (2025)

Testing the Deliteralization Hypothesis in Human and Machine Translation
by: Marmonier, Malik, et al.
Published: (2026)

Explicit Learning and the LLM in Machine Translation
by: Marmonier, Malik, et al.
Published: (2025)

Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
by: Marmonier, Malik, et al.
Published: (2026)

Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)

Towards Zero-Shot Multimodal Machine Translation
by: Futeral, Matthieu, et al.
Published: (2024)

Disentangling meaning from language in LLM-based machine translation
by: Lasnier, Théo, et al.
Published: (2026)

When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
by: Bafna, Niyati, et al.
Published: (2023)

Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025)

From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
by: Futeral, Matthieu, et al.
Published: (2024)

Investigating Length Issues in Document-level Machine Translation
by: Peng, Ziqian, et al.
Published: (2024)

ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
by: Antoun, Wissam, et al.
Published: (2025)

BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025)

Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021)

Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)

Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
by: Godey, Nathan, et al.
Published: (2024)

Domain Adaptation for Japanese Sentence Embeddings with Contrastive Learning based on Synthetic Sentence Generation
by: Chen, Zihao, et al.
Published: (2025)

Refining Sentence Embedding Model through Ranking Sentences Generation with Large Language Models
by: He, Liyang, et al.
Published: (2025)

PatentEval: Understanding Errors in Patent Generation
by: Zuo, You, et al.
Published: (2024)

How Should We Model the Probability of a Language?
by: Dent, Rasul, et al.
Published: (2026)

KréyoLID From Language Identification Towards Language Mining
by: Dent, Rasul, et al.
Published: (2025)

On the Scaling Laws of Geographical Representation in Language Models
by: Godey, Nathan, et al.
Published: (2024)

Space Decomposition for Sentence Embedding
by: Ponwitayarat, Wuttikorn, et al.
Published: (2024)

Text Simplification with Sentence Embeddings
by: Shardlow, Matthew
Published: (2025)

Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
by: Clérice, Thibault, et al.
Published: (2026)

Language-Switching Triggers Take a Latent Detour Through Language Models
by: Kulumba, Francis, et al.
Published: (2026)

RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
by: Asl, Javad Rafiei, et al.
Published: (2024)

Simple Techniques for Enhancing Sentence Embeddings in Generative Language Models
by: Zhang, Bowen, et al.
Published: (2024)

Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)

Set-Theoretic Compositionality of Sentence Embeddings
by: Bansal, Naman, et al.
Published: (2025)

Sentence Representations via Gaussian Embedding
by: Yoda, Shohei, et al.
Published: (2023)

Molyé: A Corpus-based Approach to Language Contact in Colonial France
by: Dent, Rasul, et al.
Published: (2024)

DoubleCCA: Improving Foundation Model Group Robustness with Random Sentence Embeddings
by: Liu, Hong, et al.
Published: (2024)

UNSEE: Unsupervised Non-contrastive Sentence Embeddings
by: Çağatan, Ömer Veysel
Published: (2024)

TransAug: Translate as Augmentation for Sentence Embeddings
by: Wang, Jue
Published: (2021)