Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
Fuente:
arXiv
Guardado en:
| Autores principales: | Cavalin, Paulo, Domingues, Pedro Henrique, Pinhanez, Claudio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
por: Gonçalves, Isabel, et al.
Publicado: (2025)
por: Gonçalves, Isabel, et al.
Publicado: (2025)
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
por: Pinhanez, Claudio, et al.
Publicado: (2025)
por: Pinhanez, Claudio, et al.
Publicado: (2025)
Computational Sentence-level Metrics Predicting Human Sentence Comprehension
por: Sun, Kun, et al.
Publicado: (2024)
por: Sun, Kun, et al.
Publicado: (2024)
Stronger Re-identification Attacks through Reasoning and Aggregation
por: Charpentier, Lucas Georges Gabriel, et al.
Publicado: (2025)
por: Charpentier, Lucas Georges Gabriel, et al.
Publicado: (2025)
Harnessing the Power of Artificial Intelligence to Vitalize Endangered Indigenous Languages: Technologies and Experiences
por: Pinhanez, Claudio, et al.
Publicado: (2024)
por: Pinhanez, Claudio, et al.
Publicado: (2024)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
por: Kodali, Prashant, et al.
Publicado: (2024)
por: Kodali, Prashant, et al.
Publicado: (2024)
A Corpus for Sentence-level Subjectivity Detection on English News Articles
por: Antici, Francesco, et al.
Publicado: (2023)
por: Antici, Francesco, et al.
Publicado: (2023)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
por: Ryan, Michael J., et al.
Publicado: (2025)
por: Ryan, Michael J., et al.
Publicado: (2025)
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction
por: Goto, Takumi, et al.
Publicado: (2024)
por: Goto, Takumi, et al.
Publicado: (2024)
Statistical Analysis of Sentence Structures through ASCII, Lexical Alignment and PCA
por: Sahdev, Abhijeet
Publicado: (2025)
por: Sahdev, Abhijeet
Publicado: (2025)
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
por: Pinhanez, Claudio, et al.
Publicado: (2024)
por: Pinhanez, Claudio, et al.
Publicado: (2024)
Rethinking Metrics for Lexical Semantic Change Detection
por: Goworek, Roksana, et al.
Publicado: (2026)
por: Goworek, Roksana, et al.
Publicado: (2026)
Unsupervised Candidate Ranking for Lexical Substitution via Holistic Sentence Semantics
por: Hu, Zhongyang, et al.
Publicado: (2025)
por: Hu, Zhongyang, et al.
Publicado: (2025)
Humans or LLMs as the Judge? A Study on Judgement Biases
por: Chen, Guiming Hardy, et al.
Publicado: (2024)
por: Chen, Guiming Hardy, et al.
Publicado: (2024)
LiMe: a Latin Corpus of Late Medieval Criminal Sentences
por: Bassani, Alessandra, et al.
Publicado: (2024)
por: Bassani, Alessandra, et al.
Publicado: (2024)
Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
por: Nguyen, Thanh-Nhi, et al.
Publicado: (2024)
por: Nguyen, Thanh-Nhi, et al.
Publicado: (2024)
Using Contextual Information for Sentence-level Morpheme Segmentation
por: Bhandari, Prabin, et al.
Publicado: (2024)
por: Bhandari, Prabin, et al.
Publicado: (2024)
Batch Aggregation: An Approach to Enhance Text Classification with Correlated Augmented Data
por: Hui, Charco, et al.
Publicado: (2025)
por: Hui, Charco, et al.
Publicado: (2025)
SkillAggregation: Reference-free LLM-Dependent Aggregation
por: Sun, Guangzhi, et al.
Publicado: (2024)
por: Sun, Guangzhi, et al.
Publicado: (2024)
To Aggregate or Not to Aggregate. That is the Question: A Case Study on Annotation Subjectivity in Span Prediction
por: Kurniawan, Kemal, et al.
Publicado: (2024)
por: Kurniawan, Kemal, et al.
Publicado: (2024)
Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics
por: Hou, Iyad Ait, et al.
Publicado: (2026)
por: Hou, Iyad Ait, et al.
Publicado: (2026)
Label Confidence Weighted Learning for Target-level Sentence Simplification
por: Qiu, Xinying, et al.
Publicado: (2024)
por: Qiu, Xinying, et al.
Publicado: (2024)
Sentence-level Media Bias Analysis with Event Relation Graph
por: Lei, Yuanyuan, et al.
Publicado: (2024)
por: Lei, Yuanyuan, et al.
Publicado: (2024)
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
por: Habash, Nizar, et al.
Publicado: (2024)
por: Habash, Nizar, et al.
Publicado: (2024)
SETUP: Sentence-level English-To-Uniform Meaning Representation Parser
por: Markle, Emma, et al.
Publicado: (2025)
por: Markle, Emma, et al.
Publicado: (2025)
Incorporating Precedents for Legal Judgement Prediction on European Court of Human Rights Cases
por: Santosh, T. Y. S. S., et al.
Publicado: (2024)
por: Santosh, T. Y. S. S., et al.
Publicado: (2024)
Direct Judgement Preference Optimization
por: Wang, Peifeng, et al.
Publicado: (2024)
por: Wang, Peifeng, et al.
Publicado: (2024)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
por: Qiu, Wenjie, et al.
Publicado: (2025)
por: Qiu, Wenjie, et al.
Publicado: (2025)
Typologically Informed Parameter Aggregation
por: Accou, Stef, et al.
Publicado: (2026)
por: Accou, Stef, et al.
Publicado: (2026)
A Semantic Distance Metric Learning approach for Lexical Semantic Change Detection
por: Aida, Taichi, et al.
Publicado: (2024)
por: Aida, Taichi, et al.
Publicado: (2024)
The Whole is Better than the Sum: Using Aggregated Demonstrations in In-Context Learning for Sequential Recommendation
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
DEPLAIN: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification
por: Stodden, Regina, et al.
Publicado: (2023)
por: Stodden, Regina, et al.
Publicado: (2023)
Human-LLM Hybrid Text Answer Aggregation for Crowd Annotations
por: Li, Jiyi
Publicado: (2024)
por: Li, Jiyi
Publicado: (2024)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
por: Zeng, Zhiyuan, et al.
Publicado: (2026)
por: Zeng, Zhiyuan, et al.
Publicado: (2026)
DALR: Dual-level Alignment Learning for Multimodal Sentence Representation Learning
por: He, Kang, et al.
Publicado: (2025)
por: He, Kang, et al.
Publicado: (2025)
ImpScore: A Learnable Metric For Quantifying The Implicitness Level of Sentence
por: Wang, Yuxin, et al.
Publicado: (2024)
por: Wang, Yuxin, et al.
Publicado: (2024)
Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
por: Zhang, Luke, et al.
Publicado: (2026)
por: Zhang, Luke, et al.
Publicado: (2026)
Ejemplares similares
-
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
por: Gonçalves, Isabel, et al.
Publicado: (2025) -
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
por: Cavalin, Paulo, et al.
Publicado: (2025) -
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025) -
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
por: Pinhanez, Claudio, et al.
Publicado: (2025) -
Computational Sentence-level Metrics Predicting Human Sentence Comprehension
por: Sun, Kun, et al.
Publicado: (2024)