The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
Fuente:
arXiv
Saved in:
| Main Authors: | Gómez-Rodríguez, Carlos, Williams, Paul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
by: Sun, Jingyi, et al.
Published: (2024)
by: Sun, Jingyi, et al.
Published: (2024)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
Nested Named Entity Recognition as Single-Pass Sequence Labeling
by: Muñoz-Ortiz, Alberto, et al.
Published: (2025)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2025)
Dancing in the syntax forest: fast, accurate and explainable sentiment analysis with SALSA
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024)
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024)
Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
by: Bao, Tong, et al.
Published: (2025)
by: Bao, Tong, et al.
Published: (2025)
Can LLMs Compute with Reasons?
by: Sandilya, Harshit, et al.
Published: (2024)
by: Sandilya, Harshit, et al.
Published: (2024)
LLMs for Legal Subsumption in German Employment Contracts
by: Wardas, Oliver, et al.
Published: (2025)
by: Wardas, Oliver, et al.
Published: (2025)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
by: Balter, Samuel G., et al.
Published: (2026)
by: Balter, Samuel G., et al.
Published: (2026)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
by: Wang, Yongjie, et al.
Published: (2025)
by: Wang, Yongjie, et al.
Published: (2025)
Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
by: Fang, Xi, et al.
Published: (2024)
by: Fang, Xi, et al.
Published: (2024)
Exploring Italian sentence embeddings properties through multi-tasking
by: Nastase, Vivi, et al.
Published: (2024)
by: Nastase, Vivi, et al.
Published: (2024)
Evaluating Pixel Language Models on Non-Standardized Languages
by: Muñoz-Ortiz, Alberto, et al.
Published: (2024)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2024)
Tracking linguistic information in transformer-based sentence embeddings through targeted sparsification
by: Nastase, Vivi, et al.
Published: (2024)
by: Nastase, Vivi, et al.
Published: (2024)
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
by: Nastase, Vivi, et al.
Published: (2024)
by: Nastase, Vivi, et al.
Published: (2024)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
by: Bayram, M. Ali, et al.
Published: (2024)
by: Bayram, M. Ali, et al.
Published: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
by: Weigang, Li, et al.
Published: (2025)
by: Weigang, Li, et al.
Published: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
by: Yao, Ben, et al.
Published: (2025)
by: Yao, Ben, et al.
Published: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
by: Tu, Songjun, et al.
Published: (2026)
by: Tu, Songjun, et al.
Published: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
by: Alzetta, Chiara, et al.
Published: (2025)
by: Alzetta, Chiara, et al.
Published: (2025)
How much do LLMs learn from negative examples?
by: Hamdan, Shadi, et al.
Published: (2025)
by: Hamdan, Shadi, et al.
Published: (2025)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
The Superalignment of Superhuman Intelligence with Large Language Models
by: Huang, Minlie, et al.
Published: (2024)
by: Huang, Minlie, et al.
Published: (2024)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
by: Seo, Yeongbin, et al.
Published: (2024)
by: Seo, Yeongbin, et al.
Published: (2024)
Experimentation in Content Moderation using RWKV
by: Yildirim, Umut, et al.
Published: (2024)
by: Yildirim, Umut, et al.
Published: (2024)
Branching Narratives: Character Decision Points Detection
by: Tikhonov, Alexey
Published: (2024)
by: Tikhonov, Alexey
Published: (2024)
On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations
by: Meng, Shiao, et al.
Published: (2024)
by: Meng, Shiao, et al.
Published: (2024)
Are there identifiable structural parts in the sentence embedding whole?
by: Nastase, Vivi, et al.
Published: (2024)
by: Nastase, Vivi, et al.
Published: (2024)
A Survey on Natural Language Counterfactual Generation
by: Wang, Yongjie, et al.
Published: (2024)
by: Wang, Yongjie, et al.
Published: (2024)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
Distilling Large Language Models for Efficient Clinical Information Extraction
by: Vedula, Karthik S., et al.
Published: (2024)
by: Vedula, Karthik S., et al.
Published: (2024)
MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-Tuning
by: Ren, Pengjie, et al.
Published: (2024)
by: Ren, Pengjie, et al.
Published: (2024)
From Brazilian Portuguese to European Portuguese
by: Sanches, João, et al.
Published: (2024)
by: Sanches, João, et al.
Published: (2024)
Measuring text summarization factuality using atomic facts entailment metrics in the context of retrieval augmented generation
by: Kriman, N. E.
Published: (2024)
by: Kriman, N. E.
Published: (2024)
Math Natural Language Inference: this should be easy!
by: de Paiva, Valeria, et al.
Published: (2025)
by: de Paiva, Valeria, et al.
Published: (2025)
Similar Items
-
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
by: Sun, Jingyi, et al.
Published: (2024) -
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023) -
Nested Named Entity Recognition as Single-Pass Sequence Labeling
by: Muñoz-Ortiz, Alberto, et al.
Published: (2025) -
Dancing in the syntax forest: fast, accurate and explainable sentiment analysis with SALSA
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024) -
Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
by: Bao, Tong, et al.
Published: (2025)