SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Chopra, Muskaan, Sparrenberg, Lorenz, Sifa, Rafet |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
A Survey on Current Trends and Recent Advances in Text Anonymization
by: Deußer, Tobias, et al.
Published: (2025)
by: Deußer, Tobias, et al.
Published: (2025)
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
by: Moustafa, Wesam, et al.
Published: (2026)
by: Moustafa, Wesam, et al.
Published: (2026)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law
by: Bashir, Ali Hamza, et al.
Published: (2026)
by: Bashir, Ali Hamza, et al.
Published: (2026)
[Vision Paper] PRObot: Enhancing Patient-Reported Outcome Measures for Diabetic Retinopathy using Chatbots and Generative AI
by: Pielka, Maren, et al.
Published: (2024)
by: Pielka, Maren, et al.
Published: (2024)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
by: Hillebrand, Lars, et al.
Published: (2024)
by: Hillebrand, Lars, et al.
Published: (2024)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
by: Berghaus, David, et al.
Published: (2025)
by: Berghaus, David, et al.
Published: (2025)
SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
by: Kazemi, Arefeh, et al.
Published: (2025)
by: Kazemi, Arefeh, et al.
Published: (2025)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
by: Schaaff, Kristina, et al.
Published: (2023)
by: Schaaff, Kristina, et al.
Published: (2023)
Curation of a Palaeohispanic Dataset for Machine Learning
by: Martínez-Fernández, Gonzalo, et al.
Published: (2026)
by: Martínez-Fernández, Gonzalo, et al.
Published: (2026)
Syn-TurnTurk: A Synthetic Dataset for Turn-Taking Prediction in Turkish Dialogues
by: Bayrak, Ahmet Tuğrul, et al.
Published: (2026)
by: Bayrak, Ahmet Tuğrul, et al.
Published: (2026)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
by: Jourdan, Fanny, et al.
Published: (2025)
by: Jourdan, Fanny, et al.
Published: (2025)
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks
by: Shliselberg, Michael, et al.
Published: (2024)
by: Shliselberg, Michael, et al.
Published: (2024)
Is continuous CoT better suited for multi-lingual reasoning?
by: Bashir, Ali Hamza, et al.
Published: (2026)
by: Bashir, Ali Hamza, et al.
Published: (2026)
TreePrompt: Leveraging Hierarchical Few-Shot Example Selection for Improved English-Persian and English-German Translation
by: Kakavand, Ramtin, et al.
Published: (2025)
by: Kakavand, Ramtin, et al.
Published: (2025)
ToxSyn: Reducing Bias in Hate Speech Detection via Synthetic Minority Data in Brazilian Portuguese
by: Brito, Iago Alves, et al.
Published: (2025)
by: Brito, Iago Alves, et al.
Published: (2025)
QueEn: A Large Language Model for Quechua-English Translation
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Rare but Severe Neural Machine Translation Errors Induced by Minimal Deletion: An Empirical Study on Chinese and English
by: Shi, Ruikang, et al.
Published: (2022)
by: Shi, Ruikang, et al.
Published: (2022)
Exploring the Potential of Machine Translation for Generating Named Entity Datasets: A Case Study between Persian and English
by: Sartipi, Amir, et al.
Published: (2023)
by: Sartipi, Amir, et al.
Published: (2023)
Should We be Pedantic About Reasoning Errors in Machine Translation?
by: Bao, Calvin, et al.
Published: (2026)
by: Bao, Calvin, et al.
Published: (2026)
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory Compliance
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
The Role of Handling Attributive Nouns in Improving Chinese-To-English Machine Translation
by: Wang, Lisa, et al.
Published: (2024)
by: Wang, Lisa, et al.
Published: (2024)
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
by: Rosenbaum, Andy, et al.
Published: (2026)
by: Rosenbaum, Andy, et al.
Published: (2026)
InstaTrans: An Instruction-Aware Translation Framework for Non-English Instruction Datasets
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
ATHAR: A High-Quality and Diverse Dataset for Classical Arabic to English Translation
by: Khalil, Mohammed, et al.
Published: (2024)
by: Khalil, Mohammed, et al.
Published: (2024)
Reasoning LLMs in the Medical Domain: A Literature Survey
by: Berger, Armin, et al.
Published: (2025)
by: Berger, Armin, et al.
Published: (2025)
Improving Non-autoregressive Machine Translation with Error Exposure and Consistency Regularization
by: Chen, Xinran, et al.
Published: (2024)
by: Chen, Xinran, et al.
Published: (2024)
Machine-Assisted Script Curation
by: Ciosici, Manuel R., et al.
Published: (2021)
by: Ciosici, Manuel R., et al.
Published: (2021)
SynCABEL: Synthetic Contextualized Augmentation for Biomedical Entity Linking
by: Remaki, Adam, et al.
Published: (2026)
by: Remaki, Adam, et al.
Published: (2026)
Significance of Chain of Thought in Gender Bias Mitigation for English-Dravidian Machine Translation
by: Prahallad, Lavanya, et al.
Published: (2024)
by: Prahallad, Lavanya, et al.
Published: (2024)
Sentiment Analysis Across Languages: Evaluation Before and After Machine Translation to English
by: Kathunia, Aekansh, et al.
Published: (2024)
by: Kathunia, Aekansh, et al.
Published: (2024)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021)
by: Gowda, Thamme, et al.
Published: (2021)
Similar Items
-
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
by: Chopra, Muskaan, et al.
Published: (2026) -
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
by: Chopra, Muskaan, et al.
Published: (2025) -
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
by: Chopra, Muskaan, et al.
Published: (2026) -
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
by: Chopra, Muskaan, et al.
Published: (2025) -
A Survey on Current Trends and Recent Advances in Text Anonymization
by: Deußer, Tobias, et al.
Published: (2025)