BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Gisserot-Boukhlef, Hippolyte, Boizard, Nicolas, Malherbe, Emmanuel, Hudelot, Céline, Colombo, Pierre |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs
por: Boizard, Nicolas, et al.
Publicado: (2026)
por: Boizard, Nicolas, et al.
Publicado: (2026)
Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2024)
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2024)
When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
por: Boizard, Nicolas, et al.
Publicado: (2025)
por: Boizard, Nicolas, et al.
Publicado: (2025)
Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2024)
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2024)
Should We Still Pretrain Encoders with Masked Language Modeling?
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2025)
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2025)
EuroBERT: Scaling Multilingual Encoders for European Languages
por: Boizard, Nicolas, et al.
Publicado: (2025)
por: Boizard, Nicolas, et al.
Publicado: (2025)
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
por: Boizard, Nicolas, et al.
Publicado: (2024)
por: Boizard, Nicolas, et al.
Publicado: (2024)
EuroLLM-22B: Technical Report
por: Ramos, Miguel Moura, et al.
Publicado: (2026)
por: Ramos, Miguel Moura, et al.
Publicado: (2026)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
por: Aswal, Darpan, et al.
Publicado: (2025)
por: Aswal, Darpan, et al.
Publicado: (2025)
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
por: Kim, Mihyeon, et al.
Publicado: (2025)
por: Kim, Mihyeon, et al.
Publicado: (2025)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
por: Lin, Wei-Hsiang, et al.
Publicado: (2025)
por: Lin, Wei-Hsiang, et al.
Publicado: (2025)
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
por: Moslonka, Charles, et al.
Publicado: (2025)
por: Moslonka, Charles, et al.
Publicado: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026)
por: Kostić, Bogdan, et al.
Publicado: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
por: Zhang, Wenbo, et al.
Publicado: (2026)
por: Zhang, Wenbo, et al.
Publicado: (2026)
Permutation-Consensus Listwise Judging for Robust Factuality Evaluation
por: Huang, Tianyi, et al.
Publicado: (2026)
por: Huang, Tianyi, et al.
Publicado: (2026)
VERT: Reliable LLM Judges for Radiology Report Evaluation
por: Bologna, Federica, et al.
Publicado: (2026)
por: Bologna, Federica, et al.
Publicado: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
por: D'Souza, Jennifer, et al.
Publicado: (2025)
por: D'Souza, Jennifer, et al.
Publicado: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
por: Thakur, Aman Singh, et al.
Publicado: (2024)
por: Thakur, Aman Singh, et al.
Publicado: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
por: Saha, Swarnadeep, et al.
Publicado: (2025)
por: Saha, Swarnadeep, et al.
Publicado: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
por: Enguehard, Joseph, et al.
Publicado: (2025)
por: Enguehard, Joseph, et al.
Publicado: (2025)
Reliable News or Propagandist News? A Neurosymbolic Model Using Genre, Topic, and Persuasion Techniques to Improve Robustness in Classification
por: Faye, Géraud, et al.
Publicado: (2026)
por: Faye, Géraud, et al.
Publicado: (2026)
Towards Robustness of Text-to-Visualization Translation against Lexical and Phrasal Variability
por: Lu, Jinwei, et al.
Publicado: (2024)
por: Lu, Jinwei, et al.
Publicado: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
por: Schwinn, Leo, et al.
Publicado: (2026)
por: Schwinn, Leo, et al.
Publicado: (2026)
A Knowledge-Enhanced Disease Diagnosis Method Based on Prompt Learning and BERT Integration
por: Zheng, Zhang
Publicado: (2024)
por: Zheng, Zhang
Publicado: (2024)
ARC-Encoder: learning compressed text representations for large language models
por: Pilchen, Hippolyte, et al.
Publicado: (2025)
por: Pilchen, Hippolyte, et al.
Publicado: (2025)
VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
por: Verga, Pat, et al.
Publicado: (2024)
por: Verga, Pat, et al.
Publicado: (2024)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
por: Shi, Zhichao, et al.
Publicado: (2025)
por: Shi, Zhichao, et al.
Publicado: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
por: Jin, Jiho, et al.
Publicado: (2026)
por: Jin, Jiho, et al.
Publicado: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
por: Chhabra, Mukul, et al.
Publicado: (2026)
por: Chhabra, Mukul, et al.
Publicado: (2026)
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
por: Kumar, Shachi H, et al.
Publicado: (2024)
por: Kumar, Shachi H, et al.
Publicado: (2024)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
por: Spiliopoulou, Evangelia, et al.
Publicado: (2025)
por: Spiliopoulou, Evangelia, et al.
Publicado: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
por: Wang, Yidong, et al.
Publicado: (2025)
por: Wang, Yidong, et al.
Publicado: (2025)
EuroLLM-9B: Technical Report
por: Martins, Pedro Henrique, et al.
Publicado: (2025)
por: Martins, Pedro Henrique, et al.
Publicado: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
Unitary Multi-Margin BERT for Robust Natural Language Processing
por: Chang, Hao-Yuan, et al.
Publicado: (2024)
por: Chang, Hao-Yuan, et al.
Publicado: (2024)
Ejemplares similares
-
BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs
por: Boizard, Nicolas, et al.
Publicado: (2026) -
Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2024) -
When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
por: Boizard, Nicolas, et al.
Publicado: (2025) -
Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2024) -
Should We Still Pretrain Encoders with Masked Language Modeling?
por: Gisserot-Boukhlef, Hippolyte, et al.
Publicado: (2025)