Training Models on Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning
Fuente:
arXiv
Guardado en:
| Autor principal: | Kunz, Jenny |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
por: Kunz, Jenny, et al.
Publicado: (2026)
por: Kunz, Jenny, et al.
Publicado: (2026)
Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
por: Li, Yafu, et al.
Publicado: (2025)
por: Li, Yafu, et al.
Publicado: (2025)
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
por: Liu, Yikang, et al.
Publicado: (2025)
por: Liu, Yikang, et al.
Publicado: (2025)
Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT
por: Kunz, Jenny
Publicado: (2024)
por: Kunz, Jenny
Publicado: (2024)
Pretraining Language Models Using Translationese
por: Doshi, Meet, et al.
Publicado: (2024)
por: Doshi, Meet, et al.
Publicado: (2024)
Incorporating Lexical and Syntactic Knowledge for Unsupervised Cross-Lingual Transfer
por: Zheng, Jianyu, et al.
Publicado: (2024)
por: Zheng, Jianyu, et al.
Publicado: (2024)
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers
por: Xie, Roy, et al.
Publicado: (2024)
por: Xie, Roy, et al.
Publicado: (2024)
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
por: Jang, Dongjun, et al.
Publicado: (2024)
por: Jang, Dongjun, et al.
Publicado: (2024)
How to Tune a Multilingual Encoder Model for Germanic Languages: A Study of PEFT, Full Fine-Tuning, and Language Adapters
por: Oji, Romina, et al.
Publicado: (2025)
por: Oji, Romina, et al.
Publicado: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026)
por: Kostić, Bogdan, et al.
Publicado: (2026)
Crowdsourcing Lexical Diversity
por: Khalilia, Hadi, et al.
Publicado: (2024)
por: Khalilia, Hadi, et al.
Publicado: (2024)
Preferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
por: Kunz, Jenny
Publicado: (2026)
por: Kunz, Jenny
Publicado: (2026)
A Diagnostic Benchmark for Sweden-Related Factual Knowledge
por: Kunz, Jenny
Publicado: (2025)
por: Kunz, Jenny
Publicado: (2025)
Measuring Spurious Correlation in Classification: 'Clever Hans' in Translationese
por: Borah, Angana, et al.
Publicado: (2023)
por: Borah, Angana, et al.
Publicado: (2023)
Language Model Re-rankers are Fooled by Lexical Similarities
por: Hagström, Lovisa, et al.
Publicado: (2025)
por: Hagström, Lovisa, et al.
Publicado: (2025)
A Hypothesis-Driven Framework for the Analysis of Self-Rationalising Models
por: Braun, Marc, et al.
Publicado: (2024)
por: Braun, Marc, et al.
Publicado: (2024)
Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
por: Gallagher, Daniel, et al.
Publicado: (2026)
por: Gallagher, Daniel, et al.
Publicado: (2026)
ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity
por: Jayawardena, Lasal, et al.
Publicado: (2024)
por: Jayawardena, Lasal, et al.
Publicado: (2024)
Decoding Machine Translationese in English-Chinese News: LLMs vs. NMTs
por: Kong, Delu, et al.
Publicado: (2025)
por: Kong, Delu, et al.
Publicado: (2025)
The Comparison of Translationese in Machine Translation and Human Transation in terms of Translation Relations
por: Zhou, Fan
Publicado: (2024)
por: Zhou, Fan
Publicado: (2024)
Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation
por: Wein, Shira, et al.
Publicado: (2023)
por: Wein, Shira, et al.
Publicado: (2023)
Fusing Semantic, Lexical, and Domain Perspectives for Recipe Similarity Estimation
por: Kjorvezir, Denica, et al.
Publicado: (2026)
por: Kjorvezir, Denica, et al.
Publicado: (2026)
Mitigating Translationese in Low-resource Languages: The Storyboard Approach
por: Kuwanto, Garry, et al.
Publicado: (2024)
por: Kuwanto, Garry, et al.
Publicado: (2024)
The Impact of Language Adapters in Cross-Lingual Transfer for NLU
por: Kunz, Jenny, et al.
Publicado: (2024)
por: Kunz, Jenny, et al.
Publicado: (2024)
How Lexical is Bilingual Lexicon Induction?
por: Kohli, Harsh, et al.
Publicado: (2024)
por: Kohli, Harsh, et al.
Publicado: (2024)
Vector Retrieval with Similarity and Diversity: How Hard Is It?
por: Gao, Hang, et al.
Publicado: (2024)
por: Gao, Hang, et al.
Publicado: (2024)
Code-Mixed Probes Show How Pre-Trained Models Generalise On Code-Switched Text
por: De Leon, Frances A. Laureano, et al.
Publicado: (2024)
por: De Leon, Frances A. Laureano, et al.
Publicado: (2024)
The Structural Sources of Verb Meaning Revisited: Large Language Models Display Syntactic Bootstrapping
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
por: Martinez, Richard Diehl, et al.
Publicado: (2024)
por: Martinez, Richard Diehl, et al.
Publicado: (2024)
Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation
por: Ploeger, Esther, et al.
Publicado: (2024)
por: Ploeger, Esther, et al.
Publicado: (2024)
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics
por: Lin, Fangru, et al.
Publicado: (2024)
por: Lin, Fangru, et al.
Publicado: (2024)
Modelling Child Learning and Parsing of Long-range Syntactic Dependencies
por: Mahon, Louis, et al.
Publicado: (2025)
por: Mahon, Louis, et al.
Publicado: (2025)
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
por: Karakaş, Sercan
Publicado: (2026)
por: Karakaş, Sercan
Publicado: (2026)
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models
por: Bafna, Niyati, et al.
Publicado: (2025)
por: Bafna, Niyati, et al.
Publicado: (2025)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
por: Zhang, Hongbin, et al.
Publicado: (2026)
por: Zhang, Hongbin, et al.
Publicado: (2026)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
por: Mo, Lingbo, et al.
Publicado: (2023)
por: Mo, Lingbo, et al.
Publicado: (2023)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
por: Shaib, Chantal, et al.
Publicado: (2025)
por: Shaib, Chantal, et al.
Publicado: (2025)
How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations
por: Numaya, Ikumi, et al.
Publicado: (2025)
por: Numaya, Ikumi, et al.
Publicado: (2025)
ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform
por: Lahlou, Salem
Publicado: (2026)
por: Lahlou, Salem
Publicado: (2026)
Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
por: Agarwal, Ananth, et al.
Publicado: (2025)
por: Agarwal, Ananth, et al.
Publicado: (2025)
Ejemplares similares
-
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
por: Kunz, Jenny, et al.
Publicado: (2026) -
Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
por: Li, Yafu, et al.
Publicado: (2025) -
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
por: Liu, Yikang, et al.
Publicado: (2025) -
Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT
por: Kunz, Jenny
Publicado: (2024) -
Pretraining Language Models Using Translationese
por: Doshi, Meet, et al.
Publicado: (2024)