CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
Fuente:
arXiv
Salvato in:
| Autori principali: | Cavalin, Paulo, Sanctos, Cassia, Grave, Marcelo, Pinhanez, Claudio, Primerano, Yago |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
di: Pinhanez, Claudio, et al.
Pubblicazione: (2025)
di: Pinhanez, Claudio, et al.
Pubblicazione: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
di: Cavalin, Paulo, et al.
Pubblicazione: (2025)
di: Cavalin, Paulo, et al.
Pubblicazione: (2025)
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
di: Gonçalves, Isabel, et al.
Pubblicazione: (2025)
di: Gonçalves, Isabel, et al.
Pubblicazione: (2025)
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
di: Cavalin, Paulo, et al.
Pubblicazione: (2024)
di: Cavalin, Paulo, et al.
Pubblicazione: (2024)
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
di: Pinhanez, Claudio, et al.
Pubblicazione: (2024)
di: Pinhanez, Claudio, et al.
Pubblicazione: (2024)
Harnessing the Power of Artificial Intelligence to Vitalize Endangered Indigenous Languages: Technologies and Experiences
di: Pinhanez, Claudio, et al.
Pubblicazione: (2024)
di: Pinhanez, Claudio, et al.
Pubblicazione: (2024)
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
di: Reusens, Manon, et al.
Pubblicazione: (2025)
di: Reusens, Manon, et al.
Pubblicazione: (2025)
A methodological analysis of prompt perturbations and their effect on attack success rates
di: Machado, Tiago, et al.
Pubblicazione: (2025)
di: Machado, Tiago, et al.
Pubblicazione: (2025)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
di: Chang, Chia-Hsuan, et al.
Pubblicazione: (2024)
di: Chang, Chia-Hsuan, et al.
Pubblicazione: (2024)
LLMs as Agentic Cooperative Players in Multiplayer UNO
di: Matinez, Yago Romano, et al.
Pubblicazione: (2025)
di: Matinez, Yago Romano, et al.
Pubblicazione: (2025)
Accuracy and Consistency of LLMs in the Registered Dietitian Exam: The Impact of Prompt Engineering and Knowledge Retrieval
di: Azimi, Iman, et al.
Pubblicazione: (2024)
di: Azimi, Iman, et al.
Pubblicazione: (2024)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
di: Deviyani, Athiya, et al.
Pubblicazione: (2025)
di: Deviyani, Athiya, et al.
Pubblicazione: (2025)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
di: Du, Yanrui, et al.
Pubblicazione: (2023)
di: Du, Yanrui, et al.
Pubblicazione: (2023)
Analyzing FOMC Minutes: Accuracy and Constraints of Language Models
di: Kim, Wonseong, et al.
Pubblicazione: (2023)
di: Kim, Wonseong, et al.
Pubblicazione: (2023)
To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
di: Sedova, Anastasiia, et al.
Pubblicazione: (2024)
di: Sedova, Anastasiia, et al.
Pubblicazione: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
di: Kocmi, Tom, et al.
Pubblicazione: (2024)
di: Kocmi, Tom, et al.
Pubblicazione: (2024)
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
di: Zhu, Zhihao, et al.
Pubblicazione: (2026)
di: Zhu, Zhihao, et al.
Pubblicazione: (2026)
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic
di: Robinson, Nathaniel R., et al.
Pubblicazione: (2024)
di: Robinson, Nathaniel R., et al.
Pubblicazione: (2024)
Worst-Case Input Generation for Concurrent Programs under Non-Monotone Resource Metrics
di: Pham, Long, et al.
Pubblicazione: (2023)
di: Pham, Long, et al.
Pubblicazione: (2023)
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
di: Kong, Jiawei, et al.
Pubblicazione: (2025)
di: Kong, Jiawei, et al.
Pubblicazione: (2025)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
di: Isaza-Giraldo, Andrés, et al.
Pubblicazione: (2025)
di: Isaza-Giraldo, Andrés, et al.
Pubblicazione: (2025)
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
di: Ahmed, Farhan, et al.
Pubblicazione: (2026)
di: Ahmed, Farhan, et al.
Pubblicazione: (2026)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
Attention Consistency for LLMs Explanation
di: Lan, Tian, et al.
Pubblicazione: (2025)
di: Lan, Tian, et al.
Pubblicazione: (2025)
Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet Extraction
di: Sun, Qi, et al.
Pubblicazione: (2024)
di: Sun, Qi, et al.
Pubblicazione: (2024)
LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs via Long Input Fine-Tuning
di: Mao, Yansheng, et al.
Pubblicazione: (2025)
di: Mao, Yansheng, et al.
Pubblicazione: (2025)
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo
di: Sundararajan, Barkavi, et al.
Pubblicazione: (2024)
di: Sundararajan, Barkavi, et al.
Pubblicazione: (2024)
Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
di: Tan, Hexiang, et al.
Pubblicazione: (2025)
di: Tan, Hexiang, et al.
Pubblicazione: (2025)
Comparative Experimentation of Accuracy Metrics in Automated Medical Reporting: The Case of Otitis Consultations
di: Faber, Wouter, et al.
Pubblicazione: (2023)
di: Faber, Wouter, et al.
Pubblicazione: (2023)
Moneyball with LLMs: Analyzing Tabular Summarization in Sports Narratives
di: Upadhyay, Ritam, et al.
Pubblicazione: (2025)
di: Upadhyay, Ritam, et al.
Pubblicazione: (2025)
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
di: Rao, Abhinav, et al.
Pubblicazione: (2023)
di: Rao, Abhinav, et al.
Pubblicazione: (2023)
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks
di: Pan, Eileen, et al.
Pubblicazione: (2025)
di: Pan, Eileen, et al.
Pubblicazione: (2025)
Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs
di: Muppidi, Ananth, et al.
Pubblicazione: (2025)
di: Muppidi, Ananth, et al.
Pubblicazione: (2025)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
di: Guan, Bryan, et al.
Pubblicazione: (2025)
di: Guan, Bryan, et al.
Pubblicazione: (2025)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
di: Sam, Dylan, et al.
Pubblicazione: (2025)
di: Sam, Dylan, et al.
Pubblicazione: (2025)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
di: Ghorbanpour, Faeze, et al.
Pubblicazione: (2025)
di: Ghorbanpour, Faeze, et al.
Pubblicazione: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
di: Zhao, Fufangchen, et al.
Pubblicazione: (2024)
di: Zhao, Fufangchen, et al.
Pubblicazione: (2024)
Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking
di: Aksu, Taha, et al.
Pubblicazione: (2024)
di: Aksu, Taha, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
di: Pinhanez, Claudio, et al.
Pubblicazione: (2025) -
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
di: Cavalin, Paulo, et al.
Pubblicazione: (2025) -
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
di: Gonçalves, Isabel, et al.
Pubblicazione: (2025) -
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
di: Cavalin, Paulo, et al.
Pubblicazione: (2024) -
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
di: Pinhanez, Claudio, et al.
Pubblicazione: (2024)