WRDScore: New Metric for Evaluation of Natural Language Generation Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Mussabayev, Ravil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MLLM-Microscope: Unlocking Hidden Structure Within Multimodal Large Language Models
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2026)
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2026)
Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means
von: Mussabayev, Rustam, et al.
Veröffentlicht: (2024)
von: Mussabayev, Rustam, et al.
Veröffentlicht: (2024)
Boosting K-means for Big Data by Fusing Data Streaming with Global Optimization
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2024)
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2024)
Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2023)
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2023)
High-Performance Hybrid Algorithm for Minimum Sum-of-Squares Clustering of Infinitely Tall Data
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2023)
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2023)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
von: Qi, Siya, et al.
Veröffentlicht: (2024)
von: Qi, Siya, et al.
Veröffentlicht: (2024)
Finetuning Large Language Models for Vulnerability Detection
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
Structure-Aware Code Vulnerability Analysis With Graph Neural Networks
von: Mussabayev, Ravil
Veröffentlicht: (2023)
von: Mussabayev, Ravil
Veröffentlicht: (2023)
ANLS* -- A Universal Document Processing Metric for Generative Large Language Models
von: Peer, David, et al.
Veröffentlicht: (2024)
von: Peer, David, et al.
Veröffentlicht: (2024)
Genshin: General Shield for Natural Language Processing with Large Language Models
von: Peng, Xiao, et al.
Veröffentlicht: (2024)
von: Peng, Xiao, et al.
Veröffentlicht: (2024)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
von: Shen, Ke, et al.
Veröffentlicht: (2024)
von: Shen, Ke, et al.
Veröffentlicht: (2024)
Variable Landscape Search: A Novel Metaheuristic Paradigm for Unlocking Hidden Dimensions in Global Optimization
von: Mussabayev, Rustam, et al.
Veröffentlicht: (2024)
von: Mussabayev, Rustam, et al.
Veröffentlicht: (2024)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
von: Madusanka, Tharindu, et al.
Veröffentlicht: (2025)
von: Madusanka, Tharindu, et al.
Veröffentlicht: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
ChiEngMixBench: Evaluating Large Language Models on Spontaneous and Natural Chinese-English Code-Mixed Generation
von: Yang, Qingyan, et al.
Veröffentlicht: (2026)
von: Yang, Qingyan, et al.
Veröffentlicht: (2026)
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models
von: Qiu, Ziliang, et al.
Veröffentlicht: (2025)
von: Qiu, Ziliang, et al.
Veröffentlicht: (2025)
The Evolving Landscape of Generative Large Language Models and Traditional Natural Language Processing in Medicine
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
Fake News Detection: Comparative Evaluation of BERT-like Models and Large Language Models with Generative AI-Annotated Data
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
Generating Benchmarks for Factuality Evaluation of Language Models
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
FLAME: Financial Large-Language Model Assessment and Metrics Evaluation
von: Guo, Jiayu, et al.
Veröffentlicht: (2025)
von: Guo, Jiayu, et al.
Veröffentlicht: (2025)
On the Importance and Evaluation of Narrativity in Natural Language AI Explanations
von: Cedro, Mateusz, et al.
Veröffentlicht: (2026)
von: Cedro, Mateusz, et al.
Veröffentlicht: (2026)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
von: Wu, Yulong, et al.
Veröffentlicht: (2025)
von: Wu, Yulong, et al.
Veröffentlicht: (2025)
Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules
von: Shirvani-Mahdavi, Nasim, et al.
Veröffentlicht: (2025)
von: Shirvani-Mahdavi, Nasim, et al.
Veröffentlicht: (2025)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Controlled Diversity: Length-optimized Natural Language Generation
von: Schenke, Diana Marie, et al.
Veröffentlicht: (2025)
von: Schenke, Diana Marie, et al.
Veröffentlicht: (2025)
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
von: Xu, Borui, et al.
Veröffentlicht: (2025)
von: Xu, Borui, et al.
Veröffentlicht: (2025)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
von: Ke, Pei, et al.
Veröffentlicht: (2023)
von: Ke, Pei, et al.
Veröffentlicht: (2023)
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
von: Pandey, Ashish, et al.
Veröffentlicht: (2026)
von: Pandey, Ashish, et al.
Veröffentlicht: (2026)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
von: Yang, Blair, et al.
Veröffentlicht: (2024)
von: Yang, Blair, et al.
Veröffentlicht: (2024)
Can Language Models Solve Graph Problems in Natural Language?
von: Wang, Heng, et al.
Veröffentlicht: (2023)
von: Wang, Heng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MLLM-Microscope: Unlocking Hidden Structure Within Multimodal Large Language Models
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2026) -
Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means
von: Mussabayev, Rustam, et al.
Veröffentlicht: (2024) -
Boosting K-means for Big Data by Fusing Data Streaming with Global Optimization
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2024) -
Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2023) -
High-Performance Hybrid Algorithm for Minimum Sum-of-Squares Clustering of Infinitely Tall Data
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2023)