LLMs as Span Annotators: A Comparative Study of LLMs and Humans
Fuente:
arXiv
Salvato in:
| Autori principali: | Kasner, Zdeněk, Zouhar, Vilém, Schmidtová, Patrícia, Kartáč, Ivan, Onderková, Kristýna, Plátek, Ondřej, Gkatzia, Dimitra, Mahamood, Saad, Dušek, Ondřej, Balloccu, Simone |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
factgenie: A Framework for Span-based Evaluation of Generated Texts
di: Kasner, Zdeněk, et al.
Pubblicazione: (2024)
di: Kasner, Zdeněk, et al.
Pubblicazione: (2024)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
di: Onderková, Kristýna, et al.
Pubblicazione: (2025)
di: Onderková, Kristýna, et al.
Pubblicazione: (2025)
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2024)
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2024)
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
di: Kartáč, Ivan, et al.
Pubblicazione: (2026)
di: Kartáč, Ivan, et al.
Pubblicazione: (2026)
Real-World Summarization: When Evaluation Reaches Its Limits
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2025)
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2025)
AnimatedLLM: Explaining LLMs with Interactive Visualizations
di: Kasner, Zdeněk, et al.
Pubblicazione: (2025)
di: Kasner, Zdeněk, et al.
Pubblicazione: (2025)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
di: Balloccu, Simone, et al.
Pubblicazione: (2024)
di: Balloccu, Simone, et al.
Pubblicazione: (2024)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
di: Kasner, Zdeněk, et al.
Pubblicazione: (2024)
di: Kasner, Zdeněk, et al.
Pubblicazione: (2024)
Strategies for Span Labeling with Large Language Models
di: Semin, Danil, et al.
Pubblicazione: (2026)
di: Semin, Danil, et al.
Pubblicazione: (2026)
Reasoning Gets Harder for LLMs Inside A Dialogue
di: Kartáč, Ivan, et al.
Pubblicazione: (2026)
di: Kartáč, Ivan, et al.
Pubblicazione: (2026)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
di: Kartáč, Ivan, et al.
Pubblicazione: (2025)
di: Kartáč, Ivan, et al.
Pubblicazione: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
di: Li, Karen Jia-Hui, et al.
Pubblicazione: (2025)
di: Li, Karen Jia-Hui, et al.
Pubblicazione: (2025)
A Survey of Text Style Transfer: Applications and Ethical Implications
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
di: Zouhar, Vilém, et al.
Pubblicazione: (2024)
di: Zouhar, Vilém, et al.
Pubblicazione: (2024)
Teaching LLMs at Charles University: Assignments and Activities
di: Helcl, Jindřich, et al.
Pubblicazione: (2024)
di: Helcl, Jindřich, et al.
Pubblicazione: (2024)
Multimodal Shannon Game with Images
di: Zouhar, Vilém, et al.
Pubblicazione: (2023)
di: Zouhar, Vilém, et al.
Pubblicazione: (2023)
Evaluating Optimal Reference Translations
di: Zouhar, Vilém, et al.
Pubblicazione: (2023)
di: Zouhar, Vilém, et al.
Pubblicazione: (2023)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
di: Zouhar, Vilém
Pubblicazione: (2024)
di: Zouhar, Vilém
Pubblicazione: (2024)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
di: Kocmi, Tom, et al.
Pubblicazione: (2024)
di: Kocmi, Tom, et al.
Pubblicazione: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
di: Zouhar, Vilém, et al.
Pubblicazione: (2026)
di: Zouhar, Vilém, et al.
Pubblicazione: (2026)
Text Style Transfer: An Introductory Overview
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems
di: Kumar, Nalin, et al.
Pubblicazione: (2023)
di: Kumar, Nalin, et al.
Pubblicazione: (2023)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
di: Lango, Mateusz, et al.
Pubblicazione: (2025)
di: Lango, Mateusz, et al.
Pubblicazione: (2025)
How Important is `Perfect' English for Machine Translation Prompts?
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2025)
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
AI-Assisted Human Evaluation of Machine Translation
di: Zouhar, Vilém, et al.
Pubblicazione: (2024)
di: Zouhar, Vilém, et al.
Pubblicazione: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
di: Zouhar, Vilém, et al.
Pubblicazione: (2025)
di: Zouhar, Vilém, et al.
Pubblicazione: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
di: Xu, Wenda, et al.
Pubblicazione: (2025)
di: Xu, Wenda, et al.
Pubblicazione: (2025)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
di: Sarti, Gabriele, et al.
Pubblicazione: (2025)
di: Sarti, Gabriele, et al.
Pubblicazione: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
di: Kamzela, Wiktor, et al.
Pubblicazione: (2025)
di: Kamzela, Wiktor, et al.
Pubblicazione: (2025)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
di: Warczyński, Jędrzej, et al.
Pubblicazione: (2025)
di: Warczyński, Jędrzej, et al.
Pubblicazione: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
di: Wojciechowski, Adam, et al.
Pubblicazione: (2024)
di: Wojciechowski, Adam, et al.
Pubblicazione: (2024)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
di: Bhattacharya, Sunit, et al.
Pubblicazione: (2024)
di: Bhattacharya, Sunit, et al.
Pubblicazione: (2024)
Are Large Language Models Actually Good at Text Style Transfer?
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
Sentence Embeddings as an intermediate target in end-to-end summarisation
di: Zembrzuski, Maciej, et al.
Pubblicazione: (2025)
di: Zembrzuski, Maciej, et al.
Pubblicazione: (2025)
Distributional Properties of Subword Regularization
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
Finetuning LLMs for EvaCun 2025 token prediction shared task
di: Jon, Josef, et al.
Pubblicazione: (2025)
di: Jon, Josef, et al.
Pubblicazione: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
di: Šindelář, Pavel, et al.
Pubblicazione: (2025)
di: Šindelář, Pavel, et al.
Pubblicazione: (2025)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
di: Luu, Nam, et al.
Pubblicazione: (2025)
di: Luu, Nam, et al.
Pubblicazione: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
di: Chowdhury, Sankalan Pal, et al.
Pubblicazione: (2024)
di: Chowdhury, Sankalan Pal, et al.
Pubblicazione: (2024)
Documenti analoghi
-
factgenie: A Framework for Span-based Evaluation of Generated Texts
di: Kasner, Zdeněk, et al.
Pubblicazione: (2024) -
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
di: Onderková, Kristýna, et al.
Pubblicazione: (2025) -
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2024) -
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
di: Kartáč, Ivan, et al.
Pubblicazione: (2026) -
Real-World Summarization: When Evaluation Reaches Its Limits
di: Schmidtová, Patrícia, et al.
Pubblicazione: (2025)