Recovering document annotations for sentence-level bitext
Fuente:
arXiv
Saved in:
| Main Authors: | Wicks, Rachel, Post, Matt, Koehn, Philipp |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token-level Ensembling of Models with Different Vocabularies
by: Wicks, Rachel, et al.
Published: (2025)
by: Wicks, Rachel, et al.
Published: (2025)
Escaping the sentence-level paradigm in machine translation
by: Post, Matt, et al.
Published: (2023)
by: Post, Matt, et al.
Published: (2023)
Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation
by: Liu, Ruoxi, et al.
Published: (2026)
by: Liu, Ruoxi, et al.
Published: (2026)
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
by: Meng, Chutong, et al.
Published: (2025)
by: Meng, Chutong, et al.
Published: (2025)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
by: Bafna, Niyati, et al.
Published: (2024)
by: Bafna, Niyati, et al.
Published: (2024)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
by: Proietti, Lorenzo, et al.
Published: (2026)
by: Proietti, Lorenzo, et al.
Published: (2026)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
MEXMA: Token-level objectives improve sentence representations
by: Janeiro, João Maria, et al.
Published: (2024)
by: Janeiro, João Maria, et al.
Published: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
by: Tan, Weiting, et al.
Published: (2024)
by: Tan, Weiting, et al.
Published: (2024)
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
by: Li, Tianjian, et al.
Published: (2023)
by: Li, Tianjian, et al.
Published: (2023)
POS-tagging to highlight the skeletal structure of sentences
by: Churakov, Grigorii
Published: (2024)
by: Churakov, Grigorii
Published: (2024)
Variation of sentence length across time and genre
by: Rudnicka, Karolina
Published: (2025)
by: Rudnicka, Karolina
Published: (2025)
Neural paraphrasing by automatically crawled and aligned sentence pairs
by: Globo, Achille, et al.
Published: (2024)
by: Globo, Achille, et al.
Published: (2024)
Are most sentences unique? An empirical examination of Chomskyan claims
by: Ring, Hiram
Published: (2025)
by: Ring, Hiram
Published: (2025)
X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale
by: Xu, Haoran, et al.
Published: (2024)
by: Xu, Haoran, et al.
Published: (2024)
PyMarian: Fast Neural Machine Translation and Evaluation in Python
by: Gowda, Thamme, et al.
Published: (2024)
by: Gowda, Thamme, et al.
Published: (2024)
Linguistic features for sentence difficulty prediction in ABSA
by: Chifu, Adrian-Gabriel, et al.
Published: (2024)
by: Chifu, Adrian-Gabriel, et al.
Published: (2024)
HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluation
by: Zhang, Shijie, et al.
Published: (2025)
by: Zhang, Shijie, et al.
Published: (2025)
Can LLMs capture stable human-generated sentence entropy measures?
by: Pivel-Villanueva, Estrella, et al.
Published: (2026)
by: Pivel-Villanueva, Estrella, et al.
Published: (2026)
PEACH: A sentence-aligned Parallel English-Arabic Corpus for Healthcare
by: Al-Sabbagh, Rania
Published: (2025)
by: Al-Sabbagh, Rania
Published: (2025)
Video sentence grounding with temporally global textual knowledge
by: Chen, Cai, et al.
Published: (2024)
by: Chen, Cai, et al.
Published: (2024)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
NMT-Obfuscator Attack: Ignore a sentence in translation with only one word
by: Sadrizadeh, Sahar, et al.
Published: (2024)
by: Sadrizadeh, Sahar, et al.
Published: (2024)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
by: Zhao, Xiutian, et al.
Published: (2026)
by: Zhao, Xiutian, et al.
Published: (2026)
Detection and Positive Reconstruction of Cognitive Distortion sentences: Mandarin Dataset and Evaluation
by: Lin, Shuya, et al.
Published: (2024)
by: Lin, Shuya, et al.
Published: (2024)
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
by: Kesiraju, Santosh, et al.
Published: (2026)
by: Kesiraju, Santosh, et al.
Published: (2026)
Hyperbolic sentence representations for solving Textual Entailment
by: Petrovski, Igor
Published: (2024)
by: Petrovski, Igor
Published: (2024)
Attention-aware semantic relevance predicting Chinese sentence reading
by: Sun, Kun
Published: (2024)
by: Sun, Kun
Published: (2024)
Power in Numbers: Robust reading comprehension by finetuning with four adversarial sentences per example
by: Marcus, Ariel
Published: (2024)
by: Marcus, Ariel
Published: (2024)
Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough
by: Paape, Dario, et al.
Published: (2026)
by: Paape, Dario, et al.
Published: (2026)
Understanding the effects of word-level linguistic annotations in under-resourced neural machine translation
by: Sánchez-Cartagena, Víctor M., et al.
Published: (2024)
by: Sánchez-Cartagena, Víctor M., et al.
Published: (2024)
Are there identifiable structural parts in the sentence embedding whole?
by: Nastase, Vivi, et al.
Published: (2024)
by: Nastase, Vivi, et al.
Published: (2024)
SciTaRC: Benchmarking QA on Scientific Tabular Data that Requires Language Reasoning and Complex Computation
by: Wang, Hexuan, et al.
Published: (2026)
by: Wang, Hexuan, et al.
Published: (2026)
Generating bilingual example sentences with large language models as lexicography assistants
by: Merx, Raphael, et al.
Published: (2024)
by: Merx, Raphael, et al.
Published: (2024)
LLM_annotate: A Python package for annotating and analyzing fiction characters
by: Rosenbusch, Hannes
Published: (2025)
by: Rosenbusch, Hannes
Published: (2025)
The role of inhibitory control in garden-path sentence processing: A Chinese-English bilingual perspective
by: Rao, Xiaohui, et al.
Published: (2024)
by: Rao, Xiaohui, et al.
Published: (2024)
Readers make targeted regressions to plausible errors in reanalysis of "noisy-channel garden-path" sentences
by: Clark, Thomas Hikaru, et al.
Published: (2026)
by: Clark, Thomas Hikaru, et al.
Published: (2026)
DefSent+: Improving sentence embeddings of language models by projecting definition sentences into a quasi-isotropic or isotropic vector space of unlimited dictionary entries
by: Liu, Xiaodong
Published: (2024)
by: Liu, Xiaodong
Published: (2024)
Similar Items
-
Token-level Ensembling of Models with Different Vocabularies
by: Wicks, Rachel, et al.
Published: (2025) -
Escaping the sentence-level paradigm in machine translation
by: Post, Matt, et al.
Published: (2023) -
Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation
by: Liu, Ruoxi, et al.
Published: (2026) -
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
by: Meng, Chutong, et al.
Published: (2025) -
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)