Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior
Fuente:
arXiv
Guardado en:
| Autor principal: | Nasser, Wajid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
por: Yao, Kai, et al.
Publicado: (2025)
por: Yao, Kai, et al.
Publicado: (2025)
Behavioral Fingerprints for LLM Endpoint Stability and Identity
por: Leshin, Jonah, et al.
Publicado: (2026)
por: Leshin, Jonah, et al.
Publicado: (2026)
Behavioral Fingerprinting of Large Language Models
por: Pei, Zehua, et al.
Publicado: (2025)
por: Pei, Zehua, et al.
Publicado: (2025)
Instance-level Randomization: Toward More Stable LLM Evaluations
por: Li, Yiyang, et al.
Publicado: (2025)
por: Li, Yiyang, et al.
Publicado: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026)
por: Kostić, Bogdan, et al.
Publicado: (2026)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
por: Wang, Liang, et al.
Publicado: (2026)
por: Wang, Liang, et al.
Publicado: (2026)
am-ELO: A Stable Framework for Arena-based LLM Evaluation
por: Liu, Zirui, et al.
Publicado: (2025)
por: Liu, Zirui, et al.
Publicado: (2025)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
por: Soumik, Sadman Kabir
Publicado: (2026)
por: Soumik, Sadman Kabir
Publicado: (2026)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
por: Rao, Sukrut, et al.
Publicado: (2023)
por: Rao, Sukrut, et al.
Publicado: (2023)
Evaluating LLM-Based Process Explanations under Progressive Behavioral-Input Reduction
por: van Oerle, P., et al.
Publicado: (2025)
por: van Oerle, P., et al.
Publicado: (2025)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
por: Wang, Angelina, et al.
Publicado: (2025)
por: Wang, Angelina, et al.
Publicado: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
por: Tang, Zeyu, et al.
Publicado: (2026)
por: Tang, Zeyu, et al.
Publicado: (2026)
M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games
por: Xie, Sixiong, et al.
Publicado: (2026)
por: Xie, Sixiong, et al.
Publicado: (2026)
Visual Fingerprints for LLM Generation Comparison
por: Alnouri, Amal, et al.
Publicado: (2026)
por: Alnouri, Amal, et al.
Publicado: (2026)
Behavior Alignment: A New Perspective of Evaluating LLM-based Conversational Recommender Systems
por: Yang, Dayu, et al.
Publicado: (2024)
por: Yang, Dayu, et al.
Publicado: (2024)
Attacks and Defenses Against LLM Fingerprinting
por: Kurian, Kevin, et al.
Publicado: (2025)
por: Kurian, Kevin, et al.
Publicado: (2025)
Are Robust LLM Fingerprints Adversarially Robust?
por: Nasery, Anshul, et al.
Publicado: (2025)
por: Nasery, Anshul, et al.
Publicado: (2025)
SycEval: Evaluating LLM Sycophancy
por: Fanous, Aaron, et al.
Publicado: (2025)
por: Fanous, Aaron, et al.
Publicado: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
por: Yuan, Dong, et al.
Publicado: (2024)
por: Yuan, Dong, et al.
Publicado: (2024)
LLM is Not All You Need: A Systematic Evaluation of ML vs. Foundation Models for text and image based Medical Classification
por: Raval, Meet, et al.
Publicado: (2026)
por: Raval, Meet, et al.
Publicado: (2026)
Evaluating and Understanding Scheming Propensity in LLM Agents
por: Hopman, Mia, et al.
Publicado: (2026)
por: Hopman, Mia, et al.
Publicado: (2026)
Evaluating LLM Reasoning Beyond Correctness and CoT
por: Abbasloo, Soheil
Publicado: (2025)
por: Abbasloo, Soheil
Publicado: (2025)
Towards Evaluation for Real-World LLM Unlearning
por: Miao, Ke, et al.
Publicado: (2025)
por: Miao, Ke, et al.
Publicado: (2025)
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
por: Wang, Shida, et al.
Publicado: (2025)
por: Wang, Shida, et al.
Publicado: (2025)
iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
por: Xiong, Zixun, et al.
Publicado: (2025)
por: Xiong, Zixun, et al.
Publicado: (2025)
How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation
por: Sahoo, Devanshu, et al.
Publicado: (2025)
por: Sahoo, Devanshu, et al.
Publicado: (2025)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
por: Yueh-Han, Chen, et al.
Publicado: (2025)
por: Yueh-Han, Chen, et al.
Publicado: (2025)
MEF: A Systematic Evaluation Framework for Text-to-Image Models
por: Dong, Xiaojing, et al.
Publicado: (2025)
por: Dong, Xiaojing, et al.
Publicado: (2025)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
por: Agarwal, Anisha, et al.
Publicado: (2024)
por: Agarwal, Anisha, et al.
Publicado: (2024)
Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
por: Taherinezhad, Fatemeh, et al.
Publicado: (2025)
por: Taherinezhad, Fatemeh, et al.
Publicado: (2025)
Beyond a Single Perspective: Towards a Realistic Evaluation of Website Fingerprinting Attacks
por: Deng, Xinhao, et al.
Publicado: (2025)
por: Deng, Xinhao, et al.
Publicado: (2025)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
por: Salman, Shaeke, et al.
Publicado: (2024)
por: Salman, Shaeke, et al.
Publicado: (2024)
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
por: Haznitrama, Faiz Ghifari, et al.
Publicado: (2026)
por: Haznitrama, Faiz Ghifari, et al.
Publicado: (2026)
Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
por: Naphade, Atharv, et al.
Publicado: (2026)
por: Naphade, Atharv, et al.
Publicado: (2026)
Improving Methodologies for LLM Evaluations Across Global Languages
por: Vij, Akriti, et al.
Publicado: (2026)
por: Vij, Akriti, et al.
Publicado: (2026)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
por: Zhu, Pengyu, et al.
Publicado: (2026)
por: Zhu, Pengyu, et al.
Publicado: (2026)
Evaluation and LLM-Guided Learning of ICD Coding Rationales
por: Li, Mingyang, et al.
Publicado: (2025)
por: Li, Mingyang, et al.
Publicado: (2025)
LLM-based Evaluation Policy Extraction for Ecological Modeling
por: Cheng, Qi, et al.
Publicado: (2025)
por: Cheng, Qi, et al.
Publicado: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
por: Xie, Tinghao, et al.
Publicado: (2024)
por: Xie, Tinghao, et al.
Publicado: (2024)
Ejemplares similares
-
Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
por: Yao, Kai, et al.
Publicado: (2025) -
Behavioral Fingerprints for LLM Endpoint Stability and Identity
por: Leshin, Jonah, et al.
Publicado: (2026) -
Behavioral Fingerprinting of Large Language Models
por: Pei, Zehua, et al.
Publicado: (2025) -
Instance-level Randomization: Toward More Stable LLM Evaluations
por: Li, Yiyang, et al.
Publicado: (2025) -
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026)