Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Jin Hyun, Laminchhane, Utsawb, Farooq, Umer, Sivakumar, Uma, Kumar, Arpan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
por: Chuang, Yu-Neng, et al.
Publicado: (2025)
por: Chuang, Yu-Neng, et al.
Publicado: (2025)
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
por: Mavi, Vaibhav, et al.
Publicado: (2025)
por: Mavi, Vaibhav, et al.
Publicado: (2025)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
por: Agnimo, Yedidia, et al.
Publicado: (2026)
por: Agnimo, Yedidia, et al.
Publicado: (2026)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
por: Wang, Yinong Oliver, et al.
Publicado: (2025)
por: Wang, Yinong Oliver, et al.
Publicado: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
por: Kumar, Abhishek, et al.
Publicado: (2024)
por: Kumar, Abhishek, et al.
Publicado: (2024)
Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering
por: Varzaneh, Mostafa, et al.
Publicado: (2024)
por: Varzaneh, Mostafa, et al.
Publicado: (2024)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
por: Anugraha, David, et al.
Publicado: (2024)
por: Anugraha, David, et al.
Publicado: (2024)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
por: Wang, Wenxuan, et al.
Publicado: (2022)
por: Wang, Wenxuan, et al.
Publicado: (2022)
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
por: Gabriel, Adrian Garret, et al.
Publicado: (2024)
por: Gabriel, Adrian Garret, et al.
Publicado: (2024)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
por: Li, Zihao, et al.
Publicado: (2024)
por: Li, Zihao, et al.
Publicado: (2024)
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task
por: Fujisaki, Yuya, et al.
Publicado: (2024)
por: Fujisaki, Yuya, et al.
Publicado: (2024)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
por: Collot, Stephane, et al.
Publicado: (2025)
por: Collot, Stephane, et al.
Publicado: (2025)
Predicting Task Performance with Context-aware Scaling Laws
por: Montgomery, Kyle, et al.
Publicado: (2025)
por: Montgomery, Kyle, et al.
Publicado: (2025)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
por: Bhattacharyya, Sree, et al.
Publicado: (2026)
por: Bhattacharyya, Sree, et al.
Publicado: (2026)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
por: Kulkarni, Atharva, et al.
Publicado: (2025)
por: Kulkarni, Atharva, et al.
Publicado: (2025)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
por: Jiang, Mingjian, et al.
Publicado: (2024)
por: Jiang, Mingjian, et al.
Publicado: (2024)
Adaptive Selection of LoRA Components in Privacy-Preserving Federated Learning
por: Kim, Myoungjun, et al.
Publicado: (2026)
por: Kim, Myoungjun, et al.
Publicado: (2026)
Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task
por: Lee, Hoonick, et al.
Publicado: (2024)
por: Lee, Hoonick, et al.
Publicado: (2024)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
por: He, Yuanqin, et al.
Publicado: (2024)
por: He, Yuanqin, et al.
Publicado: (2024)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2026)
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2026)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
por: Liang, Kaiqu, et al.
Publicado: (2024)
por: Liang, Kaiqu, et al.
Publicado: (2024)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
por: He, Zhonghao, et al.
Publicado: (2025)
por: He, Zhonghao, et al.
Publicado: (2025)
Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
por: Lee, Jaebok, et al.
Publicado: (2025)
por: Lee, Jaebok, et al.
Publicado: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
por: Park, Jung In, et al.
Publicado: (2024)
por: Park, Jung In, et al.
Publicado: (2024)
Robustness as an Emergent Property of Task Performance
por: Ashury-Tahan, Shir, et al.
Publicado: (2026)
por: Ashury-Tahan, Shir, et al.
Publicado: (2026)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
por: Faulkner, Ryan, et al.
Publicado: (2026)
por: Faulkner, Ryan, et al.
Publicado: (2026)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
por: Hashimoto, Wataru, et al.
Publicado: (2024)
por: Hashimoto, Wataru, et al.
Publicado: (2024)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
por: Shamrai, Maksym, et al.
Publicado: (2025)
por: Shamrai, Maksym, et al.
Publicado: (2025)
VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric
por: Nadimi, Bardia, et al.
Publicado: (2025)
por: Nadimi, Bardia, et al.
Publicado: (2025)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
por: Mackraz, Natalie, et al.
Publicado: (2024)
por: Mackraz, Natalie, et al.
Publicado: (2024)
Translating Hanja Historical Documents to Contemporary Korean and English
por: Son, Juhee, et al.
Publicado: (2022)
por: Son, Juhee, et al.
Publicado: (2022)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
por: Nayeem, Mir Tafseer, et al.
Publicado: (2025)
por: Nayeem, Mir Tafseer, et al.
Publicado: (2025)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
por: Guo, Kevin H., et al.
Publicado: (2026)
por: Guo, Kevin H., et al.
Publicado: (2026)
A Causal Lens for Evaluating Faithfulness Metrics
por: Zaman, Kerem, et al.
Publicado: (2025)
por: Zaman, Kerem, et al.
Publicado: (2025)
Confidence Regulation Neurons in Language Models
por: Stolfo, Alessandro, et al.
Publicado: (2024)
por: Stolfo, Alessandro, et al.
Publicado: (2024)
Learning to Route LLMs with Confidence Tokens
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
Emissions and Performance Trade-off Between Small and Large Language Models
por: Garg, Anandita, et al.
Publicado: (2025)
por: Garg, Anandita, et al.
Publicado: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Ejemplares similares
-
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
por: Chuang, Yu-Neng, et al.
Publicado: (2025) -
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
por: Mavi, Vaibhav, et al.
Publicado: (2025) -
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
por: Agnimo, Yedidia, et al.
Publicado: (2026) -
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
por: Wang, Yinong Oliver, et al.
Publicado: (2025) -
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)