PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Leiter, Christoph, Eger, Steffen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Explainable Evaluation Metrics for Machine Translation
di: Leiter, Christoph, et al.
Pubblicazione: (2023)
di: Leiter, Christoph, et al.
Pubblicazione: (2023)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
di: Larionov, Daniil, et al.
Pubblicazione: (2025)
di: Larionov, Daniil, et al.
Pubblicazione: (2025)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
di: Larionov, Daniil, et al.
Pubblicazione: (2024)
di: Larionov, Daniil, et al.
Pubblicazione: (2024)
BMX: Boosting Natural Language Generation Metrics with Explainability
di: Leiter, Christoph, et al.
Pubblicazione: (2022)
di: Leiter, Christoph, et al.
Pubblicazione: (2022)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
di: Larionov, Daniil, et al.
Pubblicazione: (2025)
di: Larionov, Daniil, et al.
Pubblicazione: (2025)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
di: Belouadi, Jonas, et al.
Pubblicazione: (2022)
di: Belouadi, Jonas, et al.
Pubblicazione: (2022)
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
di: Kiefer, Lotta, et al.
Pubblicazione: (2026)
di: Kiefer, Lotta, et al.
Pubblicazione: (2026)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
di: Leiter, Christoph, et al.
Pubblicazione: (2025)
di: Leiter, Christoph, et al.
Pubblicazione: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
di: Zhang, Ran, et al.
Pubblicazione: (2024)
di: Zhang, Ran, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph
di: Nechakhin, Vladyslav, et al.
Pubblicazione: (2024)
di: Nechakhin, Vladyslav, et al.
Pubblicazione: (2024)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
di: Zhang, Ran, et al.
Pubblicazione: (2023)
di: Zhang, Ran, et al.
Pubblicazione: (2023)
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
di: Wang, Zhipin, et al.
Pubblicazione: (2026)
di: Wang, Zhipin, et al.
Pubblicazione: (2026)
Argument Summarization and its Evaluation in the Era of Large Language Models
di: Altemeyer, Moritz, et al.
Pubblicazione: (2025)
di: Altemeyer, Moritz, et al.
Pubblicazione: (2025)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
di: Greisinger, Christian, et al.
Pubblicazione: (2026)
di: Greisinger, Christian, et al.
Pubblicazione: (2026)
LLM-based multi-agent poetry generation in non-cooperative environments
di: Zhang, Ran, et al.
Pubblicazione: (2024)
di: Zhang, Ran, et al.
Pubblicazione: (2024)
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
di: Chen, Yanran, et al.
Pubblicazione: (2025)
di: Chen, Yanran, et al.
Pubblicazione: (2025)
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
di: Belouadi, Jonas, et al.
Pubblicazione: (2022)
di: Belouadi, Jonas, et al.
Pubblicazione: (2022)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
di: Leiter, Christoph, et al.
Pubblicazione: (2024)
di: Leiter, Christoph, et al.
Pubblicazione: (2024)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
di: Zhang, Leixin, et al.
Pubblicazione: (2024)
di: Zhang, Leixin, et al.
Pubblicazione: (2024)
Evaluating Diversity in Automatic Poetry Generation
di: Chen, Yanran, et al.
Pubblicazione: (2024)
di: Chen, Yanran, et al.
Pubblicazione: (2024)
LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering
di: Zhang, Ran, et al.
Pubblicazione: (2025)
di: Zhang, Ran, et al.
Pubblicazione: (2025)
Prompting LLMs: Length Control for Isometric Machine Translation
di: Javorský, Dávid, et al.
Pubblicazione: (2025)
di: Javorský, Dávid, et al.
Pubblicazione: (2025)
Is there really a Citation Age Bias in NLP?
di: Nguyen, Hoa, et al.
Pubblicazione: (2024)
di: Nguyen, Hoa, et al.
Pubblicazione: (2024)
Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models
di: Shang, Yuzhe, et al.
Pubblicazione: (2026)
di: Shang, Yuzhe, et al.
Pubblicazione: (2026)
Scaling Behavior of Machine Translation with Large Language Models under Prompt Injection Attacks
di: Sun, Zhifan, et al.
Pubblicazione: (2024)
di: Sun, Zhifan, et al.
Pubblicazione: (2024)
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
di: Cui, Menglong, et al.
Pubblicazione: (2025)
di: Cui, Menglong, et al.
Pubblicazione: (2025)
Design of an Open-Source Architecture for Neural Machine Translation
di: Lankford, Séamus, et al.
Pubblicazione: (2024)
di: Lankford, Séamus, et al.
Pubblicazione: (2024)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
di: Yari, Amir Hossein, et al.
Pubblicazione: (2025)
di: Yari, Amir Hossein, et al.
Pubblicazione: (2025)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
di: Larionov, Daniil, et al.
Pubblicazione: (2024)
di: Larionov, Daniil, et al.
Pubblicazione: (2024)
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
di: Belouadi, Jonas, et al.
Pubblicazione: (2023)
di: Belouadi, Jonas, et al.
Pubblicazione: (2023)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
di: Kunilovskaya, Maria, et al.
Pubblicazione: (2026)
di: Kunilovskaya, Maria, et al.
Pubblicazione: (2026)
Shimo Lab at "Discharge Me!": Discharge Summarization by Prompt-Driven Concatenation of Electronic Health Record Sections
di: He, Yunzhen, et al.
Pubblicazione: (2024)
di: He, Yunzhen, et al.
Pubblicazione: (2024)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
di: Huang, Xu, et al.
Pubblicazione: (2024)
di: Huang, Xu, et al.
Pubblicazione: (2024)
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
di: Belouadi, Jonas, et al.
Pubblicazione: (2024)
di: Belouadi, Jonas, et al.
Pubblicazione: (2024)
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
di: Zhang, Ran, et al.
Pubblicazione: (2026)
di: Zhang, Ran, et al.
Pubblicazione: (2026)
Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality
di: Haq, Sami Ul, et al.
Pubblicazione: (2025)
di: Haq, Sami Ul, et al.
Pubblicazione: (2025)
Open Machine Translation for Esperanto
di: de Gibert, Ona, et al.
Pubblicazione: (2026)
di: de Gibert, Ona, et al.
Pubblicazione: (2026)
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
di: Jonker, Richard A. A., et al.
Pubblicazione: (2026)
di: Jonker, Richard A. A., et al.
Pubblicazione: (2026)
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
di: Herserant, Tanguy, et al.
Pubblicazione: (2025)
di: Herserant, Tanguy, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Explainable Evaluation Metrics for Machine Translation
di: Leiter, Christoph, et al.
Pubblicazione: (2023) -
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
di: Larionov, Daniil, et al.
Pubblicazione: (2025) -
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
di: Larionov, Daniil, et al.
Pubblicazione: (2024) -
BMX: Boosting Natural Language Generation Metrics with Explainability
di: Leiter, Christoph, et al.
Pubblicazione: (2022) -
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
di: Larionov, Daniil, et al.
Pubblicazione: (2025)