Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior
Fuente:
arXiv
Salvato in:
| Autore principale: | Nasser, Wajid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
di: Yao, Kai, et al.
Pubblicazione: (2025)
di: Yao, Kai, et al.
Pubblicazione: (2025)
Behavioral Fingerprints for LLM Endpoint Stability and Identity
di: Leshin, Jonah, et al.
Pubblicazione: (2026)
di: Leshin, Jonah, et al.
Pubblicazione: (2026)
Behavioral Fingerprinting of Large Language Models
di: Pei, Zehua, et al.
Pubblicazione: (2025)
di: Pei, Zehua, et al.
Pubblicazione: (2025)
Instance-level Randomization: Toward More Stable LLM Evaluations
di: Li, Yiyang, et al.
Pubblicazione: (2025)
di: Li, Yiyang, et al.
Pubblicazione: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
di: Kostić, Bogdan, et al.
Pubblicazione: (2026)
di: Kostić, Bogdan, et al.
Pubblicazione: (2026)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
di: Wang, Liang, et al.
Pubblicazione: (2026)
di: Wang, Liang, et al.
Pubblicazione: (2026)
am-ELO: A Stable Framework for Arena-based LLM Evaluation
di: Liu, Zirui, et al.
Pubblicazione: (2025)
di: Liu, Zirui, et al.
Pubblicazione: (2025)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
di: Soumik, Sadman Kabir
Pubblicazione: (2026)
di: Soumik, Sadman Kabir
Pubblicazione: (2026)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
di: Rao, Sukrut, et al.
Pubblicazione: (2023)
di: Rao, Sukrut, et al.
Pubblicazione: (2023)
Evaluating LLM-Based Process Explanations under Progressive Behavioral-Input Reduction
di: van Oerle, P., et al.
Pubblicazione: (2025)
di: van Oerle, P., et al.
Pubblicazione: (2025)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games
di: Xie, Sixiong, et al.
Pubblicazione: (2026)
di: Xie, Sixiong, et al.
Pubblicazione: (2026)
Visual Fingerprints for LLM Generation Comparison
di: Alnouri, Amal, et al.
Pubblicazione: (2026)
di: Alnouri, Amal, et al.
Pubblicazione: (2026)
Behavior Alignment: A New Perspective of Evaluating LLM-based Conversational Recommender Systems
di: Yang, Dayu, et al.
Pubblicazione: (2024)
di: Yang, Dayu, et al.
Pubblicazione: (2024)
Attacks and Defenses Against LLM Fingerprinting
di: Kurian, Kevin, et al.
Pubblicazione: (2025)
di: Kurian, Kevin, et al.
Pubblicazione: (2025)
Are Robust LLM Fingerprints Adversarially Robust?
di: Nasery, Anshul, et al.
Pubblicazione: (2025)
di: Nasery, Anshul, et al.
Pubblicazione: (2025)
SycEval: Evaluating LLM Sycophancy
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
di: Yuan, Dong, et al.
Pubblicazione: (2024)
di: Yuan, Dong, et al.
Pubblicazione: (2024)
LLM is Not All You Need: A Systematic Evaluation of ML vs. Foundation Models for text and image based Medical Classification
di: Raval, Meet, et al.
Pubblicazione: (2026)
di: Raval, Meet, et al.
Pubblicazione: (2026)
Evaluating and Understanding Scheming Propensity in LLM Agents
di: Hopman, Mia, et al.
Pubblicazione: (2026)
di: Hopman, Mia, et al.
Pubblicazione: (2026)
Evaluating LLM Reasoning Beyond Correctness and CoT
di: Abbasloo, Soheil
Pubblicazione: (2025)
di: Abbasloo, Soheil
Pubblicazione: (2025)
Towards Evaluation for Real-World LLM Unlearning
di: Miao, Ke, et al.
Pubblicazione: (2025)
di: Miao, Ke, et al.
Pubblicazione: (2025)
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
di: Wang, Shida, et al.
Pubblicazione: (2025)
di: Wang, Shida, et al.
Pubblicazione: (2025)
iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
di: Xiong, Zixun, et al.
Pubblicazione: (2025)
di: Xiong, Zixun, et al.
Pubblicazione: (2025)
How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation
di: Sahoo, Devanshu, et al.
Pubblicazione: (2025)
di: Sahoo, Devanshu, et al.
Pubblicazione: (2025)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
di: Yueh-Han, Chen, et al.
Pubblicazione: (2025)
di: Yueh-Han, Chen, et al.
Pubblicazione: (2025)
MEF: A Systematic Evaluation Framework for Text-to-Image Models
di: Dong, Xiaojing, et al.
Pubblicazione: (2025)
di: Dong, Xiaojing, et al.
Pubblicazione: (2025)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
di: Agarwal, Anisha, et al.
Pubblicazione: (2024)
di: Agarwal, Anisha, et al.
Pubblicazione: (2024)
Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
di: Taherinezhad, Fatemeh, et al.
Pubblicazione: (2025)
di: Taherinezhad, Fatemeh, et al.
Pubblicazione: (2025)
Beyond a Single Perspective: Towards a Realistic Evaluation of Website Fingerprinting Attacks
di: Deng, Xinhao, et al.
Pubblicazione: (2025)
di: Deng, Xinhao, et al.
Pubblicazione: (2025)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
di: Haznitrama, Faiz Ghifari, et al.
Pubblicazione: (2026)
di: Haznitrama, Faiz Ghifari, et al.
Pubblicazione: (2026)
Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
di: Naphade, Atharv, et al.
Pubblicazione: (2026)
di: Naphade, Atharv, et al.
Pubblicazione: (2026)
Improving Methodologies for LLM Evaluations Across Global Languages
di: Vij, Akriti, et al.
Pubblicazione: (2026)
di: Vij, Akriti, et al.
Pubblicazione: (2026)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
di: Zhu, Pengyu, et al.
Pubblicazione: (2026)
di: Zhu, Pengyu, et al.
Pubblicazione: (2026)
Evaluation and LLM-Guided Learning of ICD Coding Rationales
di: Li, Mingyang, et al.
Pubblicazione: (2025)
di: Li, Mingyang, et al.
Pubblicazione: (2025)
LLM-based Evaluation Policy Extraction for Ecological Modeling
di: Cheng, Qi, et al.
Pubblicazione: (2025)
di: Cheng, Qi, et al.
Pubblicazione: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
di: Xie, Tinghao, et al.
Pubblicazione: (2024)
di: Xie, Tinghao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
di: Yao, Kai, et al.
Pubblicazione: (2025) -
Behavioral Fingerprints for LLM Endpoint Stability and Identity
di: Leshin, Jonah, et al.
Pubblicazione: (2026) -
Behavioral Fingerprinting of Large Language Models
di: Pei, Zehua, et al.
Pubblicazione: (2025) -
Instance-level Randomization: Toward More Stable LLM Evaluations
di: Li, Yiyang, et al.
Pubblicazione: (2025) -
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
di: Kostić, Bogdan, et al.
Pubblicazione: (2026)