MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kranti, Chalamalasetti, Vajjala, Sowmya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
IndicGEC: Powerful Models, or a Measurement Mirage?
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
The Problem with Safety Classification is not just the Models
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Annotation Errors and NER: A Study with OntoNotes 5.0
by: Bernier-Colborne, Gabriel, et al.
Published: (2024)
by: Bernier-Colborne, Gabriel, et al.
Published: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026)
by: Kranti, Chalamalasetti, et al.
Published: (2026)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Text Classification in the LLM Era -- Where do we stand?
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
by: Kamath, Gaurav, et al.
Published: (2025)
by: Kamath, Gaurav, et al.
Published: (2025)
Dravidian language family through Universal Dependencies lens
by: Rama, Taraka, et al.
Published: (2024)
by: Rama, Taraka, et al.
Published: (2024)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
by: Kranti, Chalamalasetti
Published: (2025)
by: Kranti, Chalamalasetti
Published: (2025)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025)
by: Schlangen, David, et al.
Published: (2025)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
by: Beyer, Anne, et al.
Published: (2024)
by: Beyer, Anne, et al.
Published: (2024)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
Scope Ambiguities in Large Language Models
by: Kamath, Gaurav, et al.
Published: (2024)
by: Kamath, Gaurav, et al.
Published: (2024)
ARGS: Alignment as Reward-Guided Search
by: Khanov, Maxim, et al.
Published: (2024)
by: Khanov, Maxim, et al.
Published: (2024)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs
by: Parupudi, V. S. Raghu
Published: (2025)
by: Parupudi, V. S. Raghu
Published: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026)
by: Patel, Arkil, et al.
Published: (2026)
Understanding Literary Texts by LLMs: A Case Study of Ancient Chinese Poetry
by: Zhao, Cheng, et al.
Published: (2024)
by: Zhao, Cheng, et al.
Published: (2024)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
by: Deviyani, Athiya, et al.
Published: (2025)
by: Deviyani, Athiya, et al.
Published: (2025)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)
by: Kartáč, Ivan, et al.
Published: (2025)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
by: Tan, Nelvin, et al.
Published: (2026)
by: Tan, Nelvin, et al.
Published: (2026)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
by: Kommers, Cody, et al.
Published: (2025)
by: Kommers, Cody, et al.
Published: (2025)
MetRex: A Benchmark for Verilog Code Metric Reasoning Using LLMs
by: Abdelatty, Manar, et al.
Published: (2024)
by: Abdelatty, Manar, et al.
Published: (2024)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
by: Cao, Yixin, et al.
Published: (2025)
by: Cao, Yixin, et al.
Published: (2025)
Code LLMs: A Taxonomy-based Survey
by: Raihan, Nishat, et al.
Published: (2024)
by: Raihan, Nishat, et al.
Published: (2024)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
On the Robust Approximation of ASR Metrics
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs
by: Qiao, Yitong, et al.
Published: (2026)
by: Qiao, Yitong, et al.
Published: (2026)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Similar Items
-
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
IndicGEC: Powerful Models, or a Measurement Mirage?
by: Vajjala, Sowmya
Published: (2025) -
The Problem with Safety Classification is not just the Models
by: Vajjala, Sowmya
Published: (2025)