Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kasner, Zdeněk, Dušek, Ondřej |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnimatedLLM: Explaining LLMs with Interactive Visualizations
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
by: Onderková, Kristýna, et al.
Published: (2025)
by: Onderková, Kristýna, et al.
Published: (2025)
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
Strategies for Span Labeling with Large Language Models
by: Semin, Danil, et al.
Published: (2026)
by: Semin, Danil, et al.
Published: (2026)
A Survey of Text Style Transfer: Applications and Ethical Implications
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Teaching LLMs at Charles University: Assignments and Activities
by: Helcl, Jindřich, et al.
Published: (2024)
by: Helcl, Jindřich, et al.
Published: (2024)
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)
by: Kartáč, Ivan, et al.
Published: (2025)
Text Style Transfer: An Introductory Overview
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
by: Lango, Mateusz, et al.
Published: (2025)
by: Lango, Mateusz, et al.
Published: (2025)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
by: Warczyński, Jędrzej, et al.
Published: (2025)
by: Warczyński, Jędrzej, et al.
Published: (2025)
Reasoning Gets Harder for LLMs Inside A Dialogue
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
Are Large Language Models Actually Good at Text Style Transfer?
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
by: Balloccu, Simone, et al.
Published: (2024)
by: Balloccu, Simone, et al.
Published: (2024)
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems
by: Kumar, Nalin, et al.
Published: (2023)
by: Kumar, Nalin, et al.
Published: (2023)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
by: Kamzela, Wiktor, et al.
Published: (2025)
by: Kamzela, Wiktor, et al.
Published: (2025)
Real-World Summarization: When Evaluation Reaches Its Limits
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Findings of the Fourth Shared Task on Multilingual Coreference Resolution: Can LLMs Dethrone Traditional Approaches?
by: Novák, Michal, et al.
Published: (2025)
by: Novák, Michal, et al.
Published: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024)
by: Wojciechowski, Adam, et al.
Published: (2024)
Text Detoxification as Style Transfer in English and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Multilingual Text Style Transfer: Datasets & Models for Indian Languages
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
by: Schmidtová, Patrícia, et al.
Published: (2024)
by: Schmidtová, Patrícia, et al.
Published: (2024)
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks
by: Pan, Eileen, et al.
Published: (2025)
by: Pan, Eileen, et al.
Published: (2025)
Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification
by: Azqueta-Gavaldón, Andres, et al.
Published: (2025)
by: Azqueta-Gavaldón, Andres, et al.
Published: (2025)
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts
by: Li, Zehan, et al.
Published: (2025)
by: Li, Zehan, et al.
Published: (2025)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
by: Zhou, Wenrui, et al.
Published: (2025)
by: Zhou, Wenrui, et al.
Published: (2025)
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
by: Elizabeth, Michelle, et al.
Published: (2024)
by: Elizabeth, Michelle, et al.
Published: (2024)
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
by: Carlsson, Fredrik, et al.
Published: (2024)
by: Carlsson, Fredrik, et al.
Published: (2024)
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
by: Minixhofer, Christoph, et al.
Published: (2025)
by: Minixhofer, Christoph, et al.
Published: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
by: Bhattacharya, Sunit, et al.
Published: (2024)
by: Bhattacharya, Sunit, et al.
Published: (2024)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
TagRouter: Learning Route to LLMs through Tags for Open-Domain Text Generation Tasks
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Finetuning LLMs for EvaCun 2025 token prediction shared task
by: Jon, Josef, et al.
Published: (2025)
by: Jon, Josef, et al.
Published: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
Similar Items
-
AnimatedLLM: Explaining LLMs with Interactive Visualizations
by: Kasner, Zdeněk, et al.
Published: (2025) -
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
by: Onderková, Kristýna, et al.
Published: (2025) -
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024) -
Strategies for Span Labeling with Large Language Models
by: Semin, Danil, et al.
Published: (2026) -
A Survey of Text Style Transfer: Applications and Ethical Implications
by: Mukherjee, Sourabrata, et al.
Published: (2024)