Saved in:
| Main Author: | Nasser, Wajid |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.05114 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
by: Yao, Kai, et al.
Published: (2025)
by: Yao, Kai, et al.
Published: (2025)
Behavioral Fingerprints for LLM Endpoint Stability and Identity
by: Leshin, Jonah, et al.
Published: (2026)
by: Leshin, Jonah, et al.
Published: (2026)
Behavioral Fingerprinting of Large Language Models
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
Instance-level Randomization: Toward More Stable LLM Evaluations
by: Li, Yiyang, et al.
Published: (2025)
by: Li, Yiyang, et al.
Published: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
by: Kostić, Bogdan, et al.
Published: (2026)
by: Kostić, Bogdan, et al.
Published: (2026)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
by: Wang, Liang, et al.
Published: (2026)
by: Wang, Liang, et al.
Published: (2026)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
am-ELO: A Stable Framework for Arena-based LLM Evaluation
by: Liu, Zirui, et al.
Published: (2025)
by: Liu, Zirui, et al.
Published: (2025)
Visual Fingerprints for LLM Generation Comparison
by: Alnouri, Amal, et al.
Published: (2026)
by: Alnouri, Amal, et al.
Published: (2026)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
by: Soumik, Sadman Kabir
Published: (2026)
by: Soumik, Sadman Kabir
Published: (2026)
Evaluating LLM-Based Process Explanations under Progressive Behavioral-Input Reduction
by: van Oerle, P., et al.
Published: (2025)
by: van Oerle, P., et al.
Published: (2025)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games
by: Xie, Sixiong, et al.
Published: (2026)
by: Xie, Sixiong, et al.
Published: (2026)
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
Behavior Alignment: A New Perspective of Evaluating LLM-based Conversational Recommender Systems
by: Yang, Dayu, et al.
Published: (2024)
by: Yang, Dayu, et al.
Published: (2024)
SycEval: Evaluating LLM Sycophancy
by: Fanous, Aaron, et al.
Published: (2025)
by: Fanous, Aaron, et al.
Published: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
by: Yuan, Dong, et al.
Published: (2024)
by: Yuan, Dong, et al.
Published: (2024)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
by: Wang, Shida, et al.
Published: (2025)
by: Wang, Shida, et al.
Published: (2025)
iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
by: Xiong, Zixun, et al.
Published: (2025)
by: Xiong, Zixun, et al.
Published: (2025)
Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
by: Taherinezhad, Fatemeh, et al.
Published: (2025)
by: Taherinezhad, Fatemeh, et al.
Published: (2025)
Beyond a Single Perspective: Towards a Realistic Evaluation of Website Fingerprinting Attacks
by: Deng, Xinhao, et al.
Published: (2025)
by: Deng, Xinhao, et al.
Published: (2025)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation
by: Sahoo, Devanshu, et al.
Published: (2025)
by: Sahoo, Devanshu, et al.
Published: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
by: Chen, Sixu, et al.
Published: (2026)
by: Chen, Sixu, et al.
Published: (2026)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
LLM is Not All You Need: A Systematic Evaluation of ML vs. Foundation Models for text and image based Medical Classification
by: Raval, Meet, et al.
Published: (2026)
by: Raval, Meet, et al.
Published: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Evaluating and Understanding Scheming Propensity in LLM Agents
by: Hopman, Mia, et al.
Published: (2026)
by: Hopman, Mia, et al.
Published: (2026)
Evaluating LLM Reasoning Beyond Correctness and CoT
by: Abbasloo, Soheil
Published: (2025)
by: Abbasloo, Soheil
Published: (2025)
Towards Evaluation for Real-World LLM Unlearning
by: Miao, Ke, et al.
Published: (2025)
by: Miao, Ke, et al.
Published: (2025)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
MEF: A Systematic Evaluation Framework for Text-to-Image Models
by: Dong, Xiaojing, et al.
Published: (2025)
by: Dong, Xiaojing, et al.
Published: (2025)
UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification
by: Cai, Jiacheng, et al.
Published: (2024)
by: Cai, Jiacheng, et al.
Published: (2024)
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors
by: Xu, Zhenyu, et al.
Published: (2026)
by: Xu, Zhenyu, et al.
Published: (2026)
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
by: Alkaeed, Mahdi, et al.
Published: (2026)
by: Alkaeed, Mahdi, et al.
Published: (2026)
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
by: Haznitrama, Faiz Ghifari, et al.
Published: (2026)
by: Haznitrama, Faiz Ghifari, et al.
Published: (2026)
Similar Items
-
Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
by: Yao, Kai, et al.
Published: (2025) -
Behavioral Fingerprints for LLM Endpoint Stability and Identity
by: Leshin, Jonah, et al.
Published: (2026) -
Behavioral Fingerprinting of Large Language Models
by: Pei, Zehua, et al.
Published: (2025) -
Instance-level Randomization: Toward More Stable LLM Evaluations
by: Li, Yiyang, et al.
Published: (2025) -
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
by: Kostić, Bogdan, et al.
Published: (2026)