Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
Fuente:
arXiv
Guardado en:
| Autores principales: | Chopra, Muskaan, Sparrenberg, Lorenz, Sifa, Rafet |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025)
por: Chopra, Muskaan, et al.
Publicado: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025)
por: Chopra, Muskaan, et al.
Publicado: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
por: Chopra, Muskaan, et al.
Publicado: (2026)
por: Chopra, Muskaan, et al.
Publicado: (2026)
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
por: Chopra, Muskaan, et al.
Publicado: (2025)
por: Chopra, Muskaan, et al.
Publicado: (2025)
A Survey on Current Trends and Recent Advances in Text Anonymization
por: Deußer, Tobias, et al.
Publicado: (2025)
por: Deußer, Tobias, et al.
Publicado: (2025)
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
por: Moustafa, Wesam, et al.
Publicado: (2026)
por: Moustafa, Wesam, et al.
Publicado: (2026)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
por: Khanna, Sarthak, et al.
Publicado: (2025)
por: Khanna, Sarthak, et al.
Publicado: (2025)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
por: Khanna, Sarthak, et al.
Publicado: (2025)
por: Khanna, Sarthak, et al.
Publicado: (2025)
[Vision Paper] PRObot: Enhancing Patient-Reported Outcome Measures for Diabetic Retinopathy using Chatbots and Generative AI
por: Pielka, Maren, et al.
Publicado: (2024)
por: Pielka, Maren, et al.
Publicado: (2024)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
por: Hillebrand, Lars, et al.
Publicado: (2024)
por: Hillebrand, Lars, et al.
Publicado: (2024)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
por: Hillebrand, Lars, et al.
Publicado: (2025)
por: Hillebrand, Lars, et al.
Publicado: (2025)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
por: Berghaus, David, et al.
Publicado: (2025)
por: Berghaus, David, et al.
Publicado: (2025)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
por: Gogoulou, Evangelia, et al.
Publicado: (2025)
por: Gogoulou, Evangelia, et al.
Publicado: (2025)
Is continuous CoT better suited for multi-lingual reasoning?
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
Should We be Pedantic About Reasoning Errors in Machine Translation?
por: Bao, Calvin, et al.
Publicado: (2026)
por: Bao, Calvin, et al.
Publicado: (2026)
Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory Compliance
por: Hillebrand, Lars, et al.
Publicado: (2025)
por: Hillebrand, Lars, et al.
Publicado: (2025)
Improving Non-autoregressive Machine Translation with Error Exposure and Consistency Regularization
por: Chen, Xinran, et al.
Publicado: (2024)
por: Chen, Xinran, et al.
Publicado: (2024)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024)
por: Pres, Itamar, et al.
Publicado: (2024)
Reasoning LLMs in the Medical Domain: A Literature Survey
por: Berger, Armin, et al.
Publicado: (2025)
por: Berger, Armin, et al.
Publicado: (2025)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
por: Ki, Dayeon, et al.
Publicado: (2024)
por: Ki, Dayeon, et al.
Publicado: (2024)
Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models
por: Berger, Armin, et al.
Publicado: (2025)
por: Berger, Armin, et al.
Publicado: (2025)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
por: Halder, Deepon, et al.
Publicado: (2025)
por: Halder, Deepon, et al.
Publicado: (2025)
Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
por: Zhang, Songming, et al.
Publicado: (2023)
por: Zhang, Songming, et al.
Publicado: (2023)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
por: Lyu, Boxuan, et al.
Publicado: (2025)
por: Lyu, Boxuan, et al.
Publicado: (2025)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
por: Lu, Yuyin, et al.
Publicado: (2026)
por: Lu, Yuyin, et al.
Publicado: (2026)
Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task
por: Caillaut, Gaëtan, et al.
Publicado: (2024)
por: Caillaut, Gaëtan, et al.
Publicado: (2024)
Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text
por: Velasco, Dan John, et al.
Publicado: (2025)
por: Velasco, Dan John, et al.
Publicado: (2025)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
por: Guo, Pei-Fu, et al.
Publicado: (2025)
por: Guo, Pei-Fu, et al.
Publicado: (2025)
Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
por: Zhou, Tianyi, et al.
Publicado: (2025)
por: Zhou, Tianyi, et al.
Publicado: (2025)
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
por: Zhang, Zhenliang, et al.
Publicado: (2025)
por: Zhang, Zhenliang, et al.
Publicado: (2025)
Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages
por: Mao, Zhuoyuan, et al.
Publicado: (2024)
por: Mao, Zhuoyuan, et al.
Publicado: (2024)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
por: Rios-Sialer, Ian
Publicado: (2026)
por: Rios-Sialer, Ian
Publicado: (2026)
GeoResponder: Towards Building Geospatial LLMs for Time-Critical Disaster Response
por: Zguir, Ahmed El Fekih, et al.
Publicado: (2025)
por: Zguir, Ahmed El Fekih, et al.
Publicado: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
por: Zhang, Ran, et al.
Publicado: (2024)
por: Zhang, Ran, et al.
Publicado: (2024)
Towards Privacy-Preserving Machine Translation at the Inference Stage: A New Task and Benchmark
por: Shao, Wei, et al.
Publicado: (2026)
por: Shao, Wei, et al.
Publicado: (2026)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
por: Ji, Jiaming, et al.
Publicado: (2024)
por: Ji, Jiaming, et al.
Publicado: (2024)
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark
por: Taguchi, Chihiro, et al.
Publicado: (2025)
por: Taguchi, Chihiro, et al.
Publicado: (2025)
This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs
por: Wolf, Lorenz, et al.
Publicado: (2025)
por: Wolf, Lorenz, et al.
Publicado: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
por: Piot, Paloma, et al.
Publicado: (2025)
por: Piot, Paloma, et al.
Publicado: (2025)
Ejemplares similares
-
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025) -
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025) -
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
por: Chopra, Muskaan, et al.
Publicado: (2026) -
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
por: Chopra, Muskaan, et al.
Publicado: (2025) -
A Survey on Current Trends and Recent Advances in Text Anonymization
por: Deußer, Tobias, et al.
Publicado: (2025)