Confidence and Stability of Global and Pairwise Scores in NLP Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Levtsov, Georgii, Ustalov, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reliable, Reproducible, and Really Fast Leaderboards with Evalica
by: Ustalov, Dmitry
Published: (2024)
by: Ustalov, Dmitry
Published: (2024)
Sentiment Analysis for Education with R: packages, methods and practical applications
by: Misuraca, Michelangelo, et al.
Published: (2020)
by: Misuraca, Michelangelo, et al.
Published: (2020)
Observer-Based Source Localization in Tree Infection Networks via Laplace Transforms
by: O'Connor, Kesler, et al.
Published: (2025)
by: O'Connor, Kesler, et al.
Published: (2025)
A RAG Method for Source Code Inquiry Tailored to Long-Context LLMs
by: Kamiya, Toshihiro
Published: (2024)
by: Kamiya, Toshihiro
Published: (2024)
Integration of Contextual Descriptors in Ontology Alignment for Enrichment of Semantic Correspondence
by: Manziuk, Eduard, et al.
Published: (2024)
by: Manziuk, Eduard, et al.
Published: (2024)
CoSQA+: Pioneering the Multi-Choice Code Search Benchmark with Test-Driven Agents
by: Gong, Jing, et al.
Published: (2024)
by: Gong, Jing, et al.
Published: (2024)
Homeostasis: Design and Implementation of a Self-Stabilizing Compiler
by: Nougrahiya, Aman, et al.
Published: (2021)
by: Nougrahiya, Aman, et al.
Published: (2021)
flowengineR: A Modular and Extensible Framework for Fair and Reproducible Workflow Design in R
by: Willer, Maximilian, et al.
Published: (2025)
by: Willer, Maximilian, et al.
Published: (2025)
Pairwise Comparison for Bias Identification and Quantification
by: Haak, Fabian, et al.
Published: (2025)
by: Haak, Fabian, et al.
Published: (2025)
Lightweight Transformers for Zero-Shot and Fine-Tuned Text-to-SQL Generation Using Spider
by: Seth, Chirag, et al.
Published: (2025)
by: Seth, Chirag, et al.
Published: (2025)
SIADAFIX: issue description response for adaptive program repair
by: Cao, Xin, et al.
Published: (2025)
by: Cao, Xin, et al.
Published: (2025)
Global Minima by Penalized Full-dimensional Scaling
by: de Leeuw, Jan
Published: (2024)
by: de Leeuw, Jan
Published: (2024)
Knowledge-Guided Multi-Agent Framework for Automated Requirements Development: A Vision
by: Huang, Jiangping, et al.
Published: (2025)
by: Huang, Jiangping, et al.
Published: (2025)
Commenting Higher-level Code Unit: Full Code, Reduced Code, or Hierarchical Code Summarization
by: Sun, Weisong, et al.
Published: (2025)
by: Sun, Weisong, et al.
Published: (2025)
ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization
by: Fang, Chunrong, et al.
Published: (2024)
by: Fang, Chunrong, et al.
Published: (2024)
Source Code Summarization in the Era of Large Language Models
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
Higress-RAG: A Holistic Optimization Framework for Enterprise Retrieval-Augmented Generation via Dual Hybrid Retrieval, Adaptive Routing, and CRAG
by: Lin, Weixi
Published: (2025)
by: Lin, Weixi
Published: (2025)
Efficient Solvers for SLOPE in R, Python, Julia, and C++
by: Larsson, Johan, et al.
Published: (2025)
by: Larsson, Johan, et al.
Published: (2025)
Convergence of SMACOF
by: De Leeuw, Jan
Published: (2024)
by: De Leeuw, Jan
Published: (2024)
Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model
by: Qiu, Tianhao, et al.
Published: (2026)
by: Qiu, Tianhao, et al.
Published: (2026)
LeanExplore: A search engine for Lean 4 declarations
by: Asher, Justin
Published: (2025)
by: Asher, Justin
Published: (2025)
SPARQL Generation with Entity Pre-trained GPT for KG Question Answering
by: Bustamante, Diego, et al.
Published: (2024)
by: Bustamante, Diego, et al.
Published: (2024)
Characterizing JavaScript Security Code Smells
by: Kambhampati, Vikas, et al.
Published: (2024)
by: Kambhampati, Vikas, et al.
Published: (2024)
Quantifying spike train synchrony and directionality: Measures and Applications
by: Kreuz, Thomas
Published: (2025)
by: Kreuz, Thomas
Published: (2025)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
by: Dinu, Ion George, et al.
Published: (2026)
by: Dinu, Ion George, et al.
Published: (2026)
Beyond Negation Detection: Comprehensive Assertion Detection Models for Clinical NLP
by: Kocaman, Veysel, et al.
Published: (2025)
by: Kocaman, Veysel, et al.
Published: (2025)
PairDistill: Pairwise Relevance Distillation for Dense Retrieval
by: Huang, Chao-Wei, et al.
Published: (2024)
by: Huang, Chao-Wei, et al.
Published: (2024)
Mode Estimation with Partial Feedback
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Iterative NLP Query Refinement for Enhancing Domain-Specific Information Retrieval: A Case Study in Career Services
by: Peimani, Elham, et al.
Published: (2024)
by: Peimani, Elham, et al.
Published: (2024)
Cost-Sensitive Evaluation for Binary Classifiers
by: Lombardo, Pierangelo, et al.
Published: (2025)
by: Lombardo, Pierangelo, et al.
Published: (2025)
Self-Aware Vector Embeddings for Retrieval-Augmented Generation: A Neuroscience-Inspired Framework for Temporal, Confidence-Weighted, and Relational Knowledge
by: Xu, Naizhong
Published: (2026)
by: Xu, Naizhong
Published: (2026)
Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations
by: Lemdiasova, Ekaterina, et al.
Published: (2026)
by: Lemdiasova, Ekaterina, et al.
Published: (2026)
Evaluating Perspectival Biases in Cross-Modal Retrieval
by: Saengsukhiran, Teerapol, et al.
Published: (2025)
by: Saengsukhiran, Teerapol, et al.
Published: (2025)
A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
by: Petcu, Roxana, et al.
Published: (2025)
by: Petcu, Roxana, et al.
Published: (2025)
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
by: Edge, Darren, et al.
Published: (2024)
by: Edge, Darren, et al.
Published: (2024)
Supporting Humans in Evaluating AI Summaries of Legal Depositions
by: Farzi, Naghmeh, et al.
Published: (2026)
by: Farzi, Naghmeh, et al.
Published: (2026)
grangersearch: An R Package for Exhaustive Granger Causality Testing with Tidyverse Integration
by: Korfiatis, Nikolaos
Published: (2026)
by: Korfiatis, Nikolaos
Published: (2026)
SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
by: Chelombitko, Iaroslav, et al.
Published: (2026)
by: Chelombitko, Iaroslav, et al.
Published: (2026)
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Narrative Fingerprints: Multi-Scale Author Identification via Novelty Curve Dynamics
by: Zimmerman, Fred, et al.
Published: (2026)
by: Zimmerman, Fred, et al.
Published: (2026)
Similar Items
-
Reliable, Reproducible, and Really Fast Leaderboards with Evalica
by: Ustalov, Dmitry
Published: (2024) -
Sentiment Analysis for Education with R: packages, methods and practical applications
by: Misuraca, Michelangelo, et al.
Published: (2020) -
Observer-Based Source Localization in Tree Infection Networks via Laplace Transforms
by: O'Connor, Kesler, et al.
Published: (2025) -
A RAG Method for Source Code Inquiry Tailored to Long-Context LLMs
by: Kamiya, Toshihiro
Published: (2024) -
Integration of Contextual Descriptors in Ontology Alignment for Enrichment of Semantic Correspondence
by: Manziuk, Eduard, et al.
Published: (2024)