MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kranti, Chalamalasetti, Vajjala, Sowmya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Test Set Quality in Multilingual LLM Evaluation
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
IndicGEC: Powerful Models, or a Measurement Mirage?
di: Vajjala, Sowmya
Pubblicazione: (2025)
di: Vajjala, Sowmya
Pubblicazione: (2025)
The Problem with Safety Classification is not just the Models
di: Vajjala, Sowmya
Pubblicazione: (2025)
di: Vajjala, Sowmya
Pubblicazione: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Annotation Errors and NER: A Study with OntoNotes 5.0
di: Bernier-Colborne, Gabriel, et al.
Pubblicazione: (2024)
di: Bernier-Colborne, Gabriel, et al.
Pubblicazione: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2026)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2026)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
Text Classification in the LLM Era -- Where do we stand?
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
di: Kamath, Gaurav, et al.
Pubblicazione: (2025)
di: Kamath, Gaurav, et al.
Pubblicazione: (2025)
Dravidian language family through Universal Dependencies lens
di: Rama, Taraka, et al.
Pubblicazione: (2024)
di: Rama, Taraka, et al.
Pubblicazione: (2024)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
di: Kranti, Chalamalasetti
Pubblicazione: (2025)
di: Kranti, Chalamalasetti
Pubblicazione: (2025)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
di: Schlangen, David, et al.
Pubblicazione: (2025)
di: Schlangen, David, et al.
Pubblicazione: (2025)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
di: Beyer, Anne, et al.
Pubblicazione: (2024)
di: Beyer, Anne, et al.
Pubblicazione: (2024)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
di: Alhafni, Bashar, et al.
Pubblicazione: (2024)
di: Alhafni, Bashar, et al.
Pubblicazione: (2024)
Scope Ambiguities in Large Language Models
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
ARGS: Alignment as Reward-Guided Search
di: Khanov, Maxim, et al.
Pubblicazione: (2024)
di: Khanov, Maxim, et al.
Pubblicazione: (2024)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
di: Berger, Uri, et al.
Pubblicazione: (2024)
di: Berger, Uri, et al.
Pubblicazione: (2024)
Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs
di: Parupudi, V. S. Raghu
Pubblicazione: (2025)
di: Parupudi, V. S. Raghu
Pubblicazione: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
di: Patel, Arkil, et al.
Pubblicazione: (2026)
di: Patel, Arkil, et al.
Pubblicazione: (2026)
Understanding Literary Texts by LLMs: A Case Study of Ancient Chinese Poetry
di: Zhao, Cheng, et al.
Pubblicazione: (2024)
di: Zhao, Cheng, et al.
Pubblicazione: (2024)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
di: Deviyani, Athiya, et al.
Pubblicazione: (2025)
di: Deviyani, Athiya, et al.
Pubblicazione: (2025)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
di: Kartáč, Ivan, et al.
Pubblicazione: (2025)
di: Kartáč, Ivan, et al.
Pubblicazione: (2025)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
di: Liu, Lei, et al.
Pubblicazione: (2024)
di: Liu, Lei, et al.
Pubblicazione: (2024)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
di: Cyberey, Hannah, et al.
Pubblicazione: (2024)
di: Cyberey, Hannah, et al.
Pubblicazione: (2024)
Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
di: Tan, Nelvin, et al.
Pubblicazione: (2026)
di: Tan, Nelvin, et al.
Pubblicazione: (2026)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
di: Juraska, Juraj, et al.
Pubblicazione: (2024)
di: Juraska, Juraj, et al.
Pubblicazione: (2024)
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
di: Kommers, Cody, et al.
Pubblicazione: (2025)
di: Kommers, Cody, et al.
Pubblicazione: (2025)
MetRex: A Benchmark for Verilog Code Metric Reasoning Using LLMs
di: Abdelatty, Manar, et al.
Pubblicazione: (2024)
di: Abdelatty, Manar, et al.
Pubblicazione: (2024)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
di: Cao, Yixin, et al.
Pubblicazione: (2025)
di: Cao, Yixin, et al.
Pubblicazione: (2025)
Code LLMs: A Taxonomy-based Survey
di: Raihan, Nishat, et al.
Pubblicazione: (2024)
di: Raihan, Nishat, et al.
Pubblicazione: (2024)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2025)
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2025)
On the Robust Approximation of ASR Metrics
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs
di: Qiao, Yitong, et al.
Pubblicazione: (2026)
di: Qiao, Yitong, et al.
Pubblicazione: (2026)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
di: Isaza-Giraldo, Andrés, et al.
Pubblicazione: (2025)
di: Isaza-Giraldo, Andrés, et al.
Pubblicazione: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025) -
Test Set Quality in Multilingual LLM Evaluation
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025) -
IndicGEC: Powerful Models, or a Measurement Mirage?
di: Vajjala, Sowmya
Pubblicazione: (2025) -
The Problem with Safety Classification is not just the Models
di: Vajjala, Sowmya
Pubblicazione: (2025)