Pretraining Language Models Using Translationese
Fuente:
arXiv
Saved in:
| Main Authors: | Doshi, Meet, Dabre, Raj, Bhattacharyya, Pushpak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
by: Gaikwad, Pranav, et al.
Published: (2024)
by: Gaikwad, Pranav, et al.
Published: (2024)
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
by: Ghosh, Poulami, et al.
Published: (2024)
by: Ghosh, Poulami, et al.
Published: (2024)
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024)
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024)
A Morphology-Based Investigation of Positional Encodings
by: Ghosh, Poulami, et al.
Published: (2024)
by: Ghosh, Poulami, et al.
Published: (2024)
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
by: Liu, Yikang, et al.
Published: (2025)
by: Liu, Yikang, et al.
Published: (2025)
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
by: Halder, Deepon, et al.
Published: (2026)
by: Halder, Deepon, et al.
Published: (2026)
Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
by: Yousofi, Waisullah, et al.
Published: (2024)
by: Yousofi, Waisullah, et al.
Published: (2024)
Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts
by: Oram, Sneha, et al.
Published: (2026)
by: Oram, Sneha, et al.
Published: (2026)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024)
by: Kashid, Harshvivek, et al.
Published: (2024)
StereoDetect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings
by: Shejole, Kaustubh Shivshankar, et al.
Published: (2025)
by: Shejole, Kaustubh Shivshankar, et al.
Published: (2025)
Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models
by: Jhalani, Manas, et al.
Published: (2024)
by: Jhalani, Manas, et al.
Published: (2024)
Facts-and-Feelings: Capturing both Objectivity and Subjectivity in Table-to-Text Generation
by: Dey, Tathagata, et al.
Published: (2024)
by: Dey, Tathagata, et al.
Published: (2024)
We Care: Multimodal Depression Detection and Knowledge Infused Mental Health Therapeutic Response Generation
by: Moon, Palash, et al.
Published: (2024)
by: Moon, Palash, et al.
Published: (2024)
P-ReMIS: Pragmatic Reasoning in Mental Health and a Social Implication
by: Oram, Sneha, et al.
Published: (2025)
by: Oram, Sneha, et al.
Published: (2025)
Main Predicate and Their Arguments as Explanation Signals For Intent Classification
by: Pimparkhede, Sameer, et al.
Published: (2025)
by: Pimparkhede, Sameer, et al.
Published: (2025)
Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages
by: Deoghare, Sourabh, et al.
Published: (2024)
by: Deoghare, Sourabh, et al.
Published: (2024)
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
Enhancing Food-Domain Question Answering with a Multimodal Knowledge Graph: Hybrid QA Generation and Diversity Analysis
by: B, Srihari K, et al.
Published: (2025)
by: B, Srihari K, et al.
Published: (2025)
Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation
by: Wein, Shira, et al.
Published: (2023)
by: Wein, Shira, et al.
Published: (2023)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
by: Pulipaka, Sidharth, et al.
Published: (2025)
by: Pulipaka, Sidharth, et al.
Published: (2025)
Mitigating Translationese in Low-resource Languages: The Storyboard Approach
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
by: Prasanjith, Pasunuti, et al.
Published: (2025)
by: Prasanjith, Pasunuti, et al.
Published: (2025)
Improving Text Style Transfer using Masked Diffusion Language Models with Inference-time Scaling
by: Padole, Tejomay Kishor, et al.
Published: (2025)
by: Padole, Tejomay Kishor, et al.
Published: (2025)
Influence Guided Sampling for Domain Adaptation of Text Retrievers
by: Doshi, Meet, et al.
Published: (2026)
by: Doshi, Meet, et al.
Published: (2026)
An Empirical Study of In-context Learning in LLMs for Machine Translation
by: Chitale, Pranjal A., et al.
Published: (2024)
by: Chitale, Pranjal A., et al.
Published: (2024)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics
by: Kumar, Sanjeev, et al.
Published: (2026)
by: Kumar, Sanjeev, et al.
Published: (2026)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
Ta-G-T: Subjectivity Capture in Table to Text Generation via RDF Graphs
by: Upasham, Ronak, et al.
Published: (2025)
by: Upasham, Ronak, et al.
Published: (2025)
Giving the Old a Fresh Spin: Quality Estimation-Assisted Constrained Decoding for Automatic Post-Editing
by: Deoghare, Sourabh, et al.
Published: (2025)
by: Deoghare, Sourabh, et al.
Published: (2025)
Measuring Spurious Correlation in Classification: 'Clever Hans' in Translationese
by: Borah, Angana, et al.
Published: (2023)
by: Borah, Angana, et al.
Published: (2023)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
by: Kaing, Hour, et al.
Published: (2025)
by: Kaing, Hour, et al.
Published: (2025)
Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
by: Li, Yafu, et al.
Published: (2025)
by: Li, Yafu, et al.
Published: (2025)
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
by: Kunz, Jenny, et al.
Published: (2026)
by: Kunz, Jenny, et al.
Published: (2026)
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation
by: Shejole, Kaustubh Shivshankar, et al.
Published: (2025)
by: Shejole, Kaustubh Shivshankar, et al.
Published: (2025)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
by: Kammakomati, Mehant, et al.
Published: (2024)
by: Kammakomati, Mehant, et al.
Published: (2024)
Recon, Answer, Verify: Agents in Search of Truth
by: Shukla, Satyam, et al.
Published: (2025)
by: Shukla, Satyam, et al.
Published: (2025)
$\textit{Grahak-Nyay:}$ Consumer Grievance Redressal through Large Language Models
by: Ganatra, Shrey, et al.
Published: (2025)
by: Ganatra, Shrey, et al.
Published: (2025)
Striking a Balance between Classical and Deep Learning Approaches in Natural Language Processing Pedagogy
by: Joshi, Aditya, et al.
Published: (2024)
by: Joshi, Aditya, et al.
Published: (2024)
Similar Items
-
How effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
by: Gaikwad, Pranav, et al.
Published: (2024) -
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
by: Ghosh, Poulami, et al.
Published: (2024) -
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024) -
A Morphology-Based Investigation of Positional Encodings
by: Ghosh, Poulami, et al.
Published: (2024) -
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
by: Liu, Yikang, et al.
Published: (2025)