Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique
Fuente:
arXiv
Salvato in:
| Autori principali: | Sawicki, Piotr, Grześ, Marek, Brown, Dan, Góes, Fabrício |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
di: Van Veen, Dave, et al.
Pubblicazione: (2023)
di: Van Veen, Dave, et al.
Pubblicazione: (2023)
Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
di: Das, Kunal Kingkar, et al.
Pubblicazione: (2025)
di: Das, Kunal Kingkar, et al.
Pubblicazione: (2025)
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry
di: Ma, Bolei, et al.
Pubblicazione: (2025)
di: Ma, Bolei, et al.
Pubblicazione: (2025)
Scaling Embeddings Outperforms Scaling Experts in Language Models
di: Liu, Hong, et al.
Pubblicazione: (2026)
di: Liu, Hong, et al.
Pubblicazione: (2026)
Enhancing Reasoning Skills in Small Persian Medical Language Models Can Outperform Large-Scale Data Training
di: Ghassabi, Mehrdad, et al.
Pubblicazione: (2025)
di: Ghassabi, Mehrdad, et al.
Pubblicazione: (2025)
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
di: Zou, Kaijian, et al.
Pubblicazione: (2025)
di: Zou, Kaijian, et al.
Pubblicazione: (2025)
Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets
di: Walsh, Melanie, et al.
Pubblicazione: (2024)
di: Walsh, Melanie, et al.
Pubblicazione: (2024)
Evaluating Large Language Models as Expert Annotators
di: Tseng, Yu-Min, et al.
Pubblicazione: (2025)
di: Tseng, Yu-Min, et al.
Pubblicazione: (2025)
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
di: Vishwanath, Krithik, et al.
Pubblicazione: (2025)
di: Vishwanath, Krithik, et al.
Pubblicazione: (2025)
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Bigger But Not Better: Small Neural Language Models Outperform Large Language Models in Detection of Thought Disorder
di: Li, Changye, et al.
Pubblicazione: (2025)
di: Li, Changye, et al.
Pubblicazione: (2025)
Do LLMs Agree on the Creativity Evaluation of Alternative Uses?
di: Rabeyah, Abdullah Al, et al.
Pubblicazione: (2024)
di: Rabeyah, Abdullah Al, et al.
Pubblicazione: (2024)
Can AI Outperform Human Experts in Creating Social Media Creatives?
di: Park, Eunkyung, et al.
Pubblicazione: (2024)
di: Park, Eunkyung, et al.
Pubblicazione: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
di: Marco, Guillermo, et al.
Pubblicazione: (2024)
di: Marco, Guillermo, et al.
Pubblicazione: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
di: Ludziejewski, Jan, et al.
Pubblicazione: (2025)
di: Ludziejewski, Jan, et al.
Pubblicazione: (2025)
Can Large Language Models (LLMs) Describe Pictures Like Children? A Comparative Corpus Study
di: Woloszyn, Hanna, et al.
Pubblicazione: (2025)
di: Woloszyn, Hanna, et al.
Pubblicazione: (2025)
Task-Specific Efficiency Analysis: When Small Language Models Outperform Large Language Models
di: Cao, Jinghan, et al.
Pubblicazione: (2026)
di: Cao, Jinghan, et al.
Pubblicazione: (2026)
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
di: Mikami, Yosuke, et al.
Pubblicazione: (2025)
di: Mikami, Yosuke, et al.
Pubblicazione: (2025)
Automating Expert-Level Medical Reasoning Evaluation of Large Language Models
di: Zhou, Shuang, et al.
Pubblicazione: (2025)
di: Zhou, Shuang, et al.
Pubblicazione: (2025)
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
di: Nam, Hyunji, et al.
Pubblicazione: (2025)
di: Nam, Hyunji, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
di: Nadas, Mihai, et al.
Pubblicazione: (2025)
di: Nadas, Mihai, et al.
Pubblicazione: (2025)
Deanthropomorphising NLP: Can a Language Model Be Conscious?
di: Shardlow, Matthew, et al.
Pubblicazione: (2022)
di: Shardlow, Matthew, et al.
Pubblicazione: (2022)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
di: Aly, Walid Mohamed, et al.
Pubblicazione: (2025)
di: Aly, Walid Mohamed, et al.
Pubblicazione: (2025)
LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought
di: Qi, Ruiyan, et al.
Pubblicazione: (2025)
di: Qi, Ruiyan, et al.
Pubblicazione: (2025)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
di: Liusie, Adian, et al.
Pubblicazione: (2023)
di: Liusie, Adian, et al.
Pubblicazione: (2023)
Can Large Language Models Master Complex Card Games?
di: Wang, Wei, et al.
Pubblicazione: (2025)
di: Wang, Wei, et al.
Pubblicazione: (2025)
Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Model
di: Donhauser, Niklas, et al.
Pubblicazione: (2026)
di: Donhauser, Niklas, et al.
Pubblicazione: (2026)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
Is Temperature the Creativity Parameter of Large Language Models?
di: Peeperkorn, Max, et al.
Pubblicazione: (2024)
di: Peeperkorn, Max, et al.
Pubblicazione: (2024)
Large Language Models Perform on Par with Experts Identifying Mental Health Factors in Adolescent Online Forums
di: Lorge, Isabelle, et al.
Pubblicazione: (2024)
di: Lorge, Isabelle, et al.
Pubblicazione: (2024)
A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
di: Cadeddu, Andrea, et al.
Pubblicazione: (2025)
di: Cadeddu, Andrea, et al.
Pubblicazione: (2025)
When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full Models
di: Chang, Ting-Yun, et al.
Pubblicazione: (2024)
di: Chang, Ting-Yun, et al.
Pubblicazione: (2024)
Feminist Interaction Techniques: Deterring Non-Consensual Screenshots with Interaction Techniques
di: Qiwei, Li, et al.
Pubblicazione: (2024)
di: Qiwei, Li, et al.
Pubblicazione: (2024)
Comparative Study of Domain Driven Terms Extraction Using Large Language Models
di: Chataut, Sandeep, et al.
Pubblicazione: (2024)
di: Chataut, Sandeep, et al.
Pubblicazione: (2024)
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
di: Guo, Hongcheng, et al.
Pubblicazione: (2025)
di: Guo, Hongcheng, et al.
Pubblicazione: (2025)
Evaluating the Impact of Compression Techniques on Task-Specific Performance of Large Language Models
di: Khanal, Bishwash, et al.
Pubblicazione: (2024)
di: Khanal, Bishwash, et al.
Pubblicazione: (2024)
Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications?
di: Intrator, Yotam, et al.
Pubblicazione: (2024)
di: Intrator, Yotam, et al.
Pubblicazione: (2024)
Predicting Emotion Intensity in Polish Political Texts: Comparing Supervised Models and Large Language Models in a Resource-Poor Language
di: Plisiecki, Hubert, et al.
Pubblicazione: (2024)
di: Plisiecki, Hubert, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
di: Van Veen, Dave, et al.
Pubblicazione: (2023) -
Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
di: Das, Kunal Kingkar, et al.
Pubblicazione: (2025) -
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry
di: Ma, Bolei, et al.
Pubblicazione: (2025) -
Scaling Embeddings Outperforms Scaling Experts in Language Models
di: Liu, Hong, et al.
Pubblicazione: (2026) -
Enhancing Reasoning Skills in Small Persian Medical Language Models Can Outperform Large-Scale Data Training
di: Ghassabi, Mehrdad, et al.
Pubblicazione: (2025)